mirror of
https://github.com/HeyPuter/puter.git
synced 2026-10-10 13:51:41 +00:00
docs(ai): keep model update documentation focused
This commit is contained in:
1 parent
266390d23d
commit
0afe8ccfdd
4 files changed
+2
-44
No files matched your search
@@ -7,8 +7,6 @@ The Puter.js AI feature allows you to integrate artificial intelligence capabili
|
||||
|
||||
You can use AI models from various providers to perform tasks such as chat, text-to-image, image-to-text, text-to-video, and text-to-speech conversion. And with the [User-Pays Model](/user-pays-model/), you don't have to set up your own API keys and top up credits, because users cover their own AI costs.
|
||||
|
||||
Chat models include Claude Sonnet 5.5 (`claude-sonnet-5-5`) and Mistral Large 4 (`mistral-large-4`, public preview). See [their availability, pricing, and example](/AI/chat/#claude-sonnet-55-and-mistral-large-4), or query [`puter.ai.listModels()`](/AI/listModels) for the current catalog.
|
||||
|
||||
## Features
|
||||
|
||||
<div style="overflow:hidden; margin-bottom: 30px;">
|
||||
|
||||
@@ -137,44 +137,6 @@ Output generated before the stream stops may still be billed.
|
||||
|
||||
We use different vendors for different models and try to use the best vendor available at the time of the request. Vendors currently include Alibaba Cloud, Anthropic, Azure OpenAI, DeepSeek, Google, Infron, Meta, MiniMax, Mistral, Moonshot AI, OpenAI, OpenRouter, Together AI, xAI, and Z.AI. Call [`puter.ai.listModelProviders()`](/AI/listModelProviders) for the current list, or pass `provider` in the options object to pin a request to one of them.
|
||||
|
||||
### Claude Sonnet 5.5 and Mistral Large 4
|
||||
|
||||
Both models accept text and images, support tool calling, and return normalized responses by default. Select them with `claude-sonnet-5-5` and `mistral-large-4` (`mistral-large-4-0` is also accepted). Existing `claude-sonnet` aliases select Sonnet 5.5; `mistral-large` and `mistral-large-latest` continue to select Large 3.
|
||||
|
||||
The direct-provider catalog uses these standard upstream rates in USD per million tokens, verified October 8, 2026:
|
||||
|
||||
| Model | Input | Cached input | Output | Context |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| Claude Sonnet 5.5 | $2 | $0.10 | $10 | 1M tokens |
|
||||
| Mistral Large 4 | $1.36 | $0.14 | $4.18 | 1M tokens |
|
||||
|
||||
[Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) is active on the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry. It supports 128K output tokens; 5-minute cache writes cost $2.50/M tokens and 1-hour writes cost $4/M tokens. Puter integrates it through the Claude API.
|
||||
|
||||
[Large 4](https://docs.mistral.ai/models/mistral-large-4-0) is available through Mistral's API in public preview, released October 6, 2026. Its open weights are [planned for the end of October](https://mistral.ai/news/mistral-large-4/). Mistral advertises a [two-week launch discount](https://docs.mistral.ai/resources/changelogs) of 50%: $0.68 input, $0.07 cached input, and $2.09 output per million tokens. Puter's catalog uses the standard rates above rather than this temporary promotion. Cached prompt tokens are billed separately from uncached input for streamed and non-streamed requests.
|
||||
|
||||
Provider availability depends on the deployment's configured API keys. Use [`listModels()`](/AI/listModels) to check the running instance and `GET /metering/allCosts` for its billed rates, which can include a deployment's AI cost factor.
|
||||
|
||||
```html
|
||||
<html>
|
||||
<body>
|
||||
<script src="https://js.puter.com/v2/"></script>
|
||||
<script>
|
||||
(async () => {
|
||||
for (const [model, provider] of [
|
||||
["claude-sonnet-5-5", "claude"],
|
||||
["mistral-large-4", "mistral"],
|
||||
]) {
|
||||
const response = await puter.ai.chat("Explain a rainbow in one sentence.", {
|
||||
model, provider, max_tokens: 1024,
|
||||
});
|
||||
puter.print(response.message.content);
|
||||
}
|
||||
})();
|
||||
</script>
|
||||
</body>
|
||||
</html>
|
||||
```
|
||||
|
||||
## Response Normalization
|
||||
|
||||
Most vendors respond in the OpenAI chat format, where `message.content` is a string and tool calls appear as `message.tool_calls`. Anthropic models historically respond in Anthropic's native format instead, where `message.content` is an array of content blocks such as `[{ type: "text", text: "..." }]`.
|
||||
|
||||
@@ -3,8 +3,8 @@
|
||||
<script src="https://js.puter.com/v2/"></script>
|
||||
<script>
|
||||
(async () => {
|
||||
const chat_resp = await puter.ai.chat('Tell me something I might not know about chemistry.', {model: 'claude-sonnet-4-6', stream: true });
|
||||
puter.print('<h1>Claude Sonnet 4.6:</h1>');
|
||||
const chat_resp = await puter.ai.chat('Tell me something I might not know about chemistry.', {model: 'claude-sonnet-5-5', stream: true });
|
||||
puter.print('<h1>Claude Sonnet 5.5:</h1>');
|
||||
for await ( const part of chat_resp ) puter.print(part?.text?.replaceAll('\n', '<br>'));
|
||||
})();
|
||||
</script>
|
||||
|
||||
@@ -71,8 +71,6 @@ Shared by chat, image generation, video, TTS, speech and OCR. Each interface and
|
||||
|
||||
An input over its limit is rejected before it reaches the model. OCR input limits are under [OCR](#ocr).
|
||||
|
||||
Claude Sonnet 5.5 has a 1M-token context window and a 128K output limit. Mistral Large 4's configured context and output ceilings are 1M tokens combined; its provider does not publish a separate output limit on the model card. Chat output is also bounded by the context remaining after the prompt and the caller's available credits. See [model availability and pricing](/AI/chat/#claude-sonnet-55-and-mistral-large-4).
|
||||
|
||||
The OpenAI- and Anthropic-compatible endpoints (`/puterai/openai/v1/*`, `/puterai/anthropic/v1/messages`) require a paid plan; a free account gets `402 subscription_required`. The same models are available to every account through `puter.ai.*` and `/drivers/call`, and the model catalogue endpoints are open to everyone.
|
||||
|
||||
`/puterai/anthropic/v1/messages/count_tokens` has its own budget of 120 requests per minute per user, with no concurrency limit, and is not charged.
|
||||
|
||||
Reference in new issue
Block a user