mirror of
https://github.com/HeyPuter/puter.git
synced 2026-10-10 22:01:40 +00:00
fix(ai-ocr): address review: playground examples, Textract options, inline size
Adds the Mistral and document-annotation img2txt examples to the playground, states that Textract takes no options, and drops the normalize mentions from the img2txt docs and types. The SDK now checks Mistral inline sources against the size a base64 upload can reach inside the 50 MB JSON body (36 MB) instead of the 10 MB Textract limit; 50 MB let 38-50 MB inputs through to a body-parser 413. Retired-model messages no longer cite vendor dates. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This commit is contained in:
1 parent
eef8258539
commit
d36da7522b
9 files changed
+117
-12
No files matched your search
@@ -44,7 +44,7 @@ export const OCR_MODELS: readonly OcrModel[] = [
|
||||
{
|
||||
id: 'mistral-ocr-4-1',
|
||||
provider: 'mistral',
|
||||
// Mistral retired 2503 but still answers it with the latest model.
|
||||
// Preserve Puter's old 2503 spelling by routing it to OCR 4.1.
|
||||
aliases: ['mistral-ocr-latest', 'mistral-ocr-4', 'mistral-ocr-2503'],
|
||||
pageUsageType: 'mistral-ocr:mistral-ocr-4-1:page',
|
||||
annotationUsageType: 'mistral-ocr:mistral-ocr-4-1:annotations:page',
|
||||
@@ -67,7 +67,7 @@ export const OCR_MODELS: readonly OcrModel[] = [
|
||||
/** Models the vendor no longer serves, with the reason callers see. */
|
||||
export const RETIRED_OCR_MODELS: Readonly<Record<string, string>> = {
|
||||
'mistral-ocr-2505':
|
||||
'Mistral retired it on 2026-05-31; use mistral-ocr-latest.',
|
||||
'Puter no longer supports this deprecated model; use mistral-ocr-latest.',
|
||||
};
|
||||
|
||||
export const DEFAULT_OCR_MODEL: Record<OcrProviderId, string> = {
|
||||
|
||||
@@ -43,7 +43,11 @@ Every call has the same shape; only the model name changes which service reads t
|
||||
| `mistral-ocr-4-0` | Mistral | Mistral OCR 4.0. |
|
||||
| `mistral-ocr-2512` (aliases `mistral-ocr-3`, `mistral-ocr-3-0`) | Mistral | Mistral OCR 3, at a lower per-page rate than OCR 4. |
|
||||
|
||||
`mistral-ocr-latest` is pinned to OCR 4.1 and moves to a newer model only when Puter adds it. `mistral-ocr-2503` still works and runs on OCR 4.1, because that is how Mistral now serves it. `mistral-ocr-2505` has been retired by Mistral and is rejected with `bad_request`. Per-page prices for each model are listed by the API at `GET /metering/allCosts`.
|
||||
`mistral-ocr-latest` is pinned to OCR 4.1 and moves to a newer model only when Puter adds it. Mistral has deprecated `mistral-ocr-2503`; Puter keeps that name as a compatibility alias for OCR 4.1. Puter rejects the deprecated `mistral-ocr-2505` with `bad_request`. Per-page prices for each model are listed by the API at `GET /metering/allCosts`.
|
||||
|
||||
#### AWS Textract options
|
||||
|
||||
AWS Textract takes no options beyond `model` and `provider`. It reads a single-page document as a whole, so there is nothing to select or tune; the Mistral options below are ignored.
|
||||
|
||||
#### Mistral options
|
||||
|
||||
@@ -73,7 +77,7 @@ For more details about each option, see the [Mistral OCR documentation](https://
|
||||
| AWS Textract | JPEG, PNG, TIFF, and **single-page** PDF | 10 MB |
|
||||
| Mistral | PDF (up to 1,000 pages), images (JPEG, PNG, AVIF, TIFF, GIF, HEIC, BMP, WebP), and documents such as DOCX, PPTX, XLSX, EPUB and RTF | 50 MB |
|
||||
|
||||
`File`, `Blob` and data URI inputs are limited to 10 MB by the SDK before upload, whichever model is used. URLs and Puter paths are limited to the model's maximum. Larger inputs are rejected with `storage_limit_reached`. Textract rejects multi-page PDFs and other formats with `bad_request`; use a Mistral model for those.
|
||||
The SDK checks the decoded size of data URI inputs against the selected model's limit before upload: 10 MB for Textract and 36 MB when a Mistral model or provider is specified (a data URI is a third larger than the file, and a request body is capped at 50 MB). URLs and Puter paths are limited by the backend. When neither model nor provider is specified, the SDK checks against 10 MB; explicitly select Mistral for larger inline inputs. SDK size failures use `input_too_large`; backend size failures use `storage_limit_reached`. Textract rejects multi-page PDFs and other formats with `bad_request`; use a Mistral model for those.
|
||||
|
||||
## Return value
|
||||
|
||||
@@ -82,7 +86,7 @@ A `Promise` that resolves to a string.
|
||||
- By default the string is the recognized text, one line per line: plain text from AWS Textract, Markdown from Mistral.
|
||||
- When `documentAnnotationFormat` is set, the string is the document annotation: JSON that follows your schema.
|
||||
|
||||
The return shape is the same for every model. That means the `normalize` option and `puter.ai.normalize` described for [`chat()`](/AI/chat#response-normalization) have no effect on `img2txt()`.
|
||||
The return shape is the same for every model.
|
||||
|
||||
## Errors
|
||||
|
||||
@@ -91,7 +95,7 @@ A rejection carries the error body as the backend sent it: `{ message, code }`.
|
||||
| Code | Meaning |
|
||||
| --- | --- |
|
||||
| `arguments_required`, `source_required` | Raised by the SDK before any request is made: the call had no arguments, or no source. |
|
||||
| `input_too_large` | Raised by the SDK before any request is made: a `File`, `Blob` or data URI input is larger than 10 MB. |
|
||||
| `input_too_large` | Raised by the SDK before any request is made: a `File`, `Blob` or data URI input exceeds the selected model's limit. |
|
||||
| `storage_limit_reached` | The input is larger than the model accepts (HTTP 413). |
|
||||
| `bad_request` | The provider or model is unknown or retired, the model does not belong to the named provider, an option is invalid, or Textract cannot read the document. |
|
||||
| `insufficient_funds` | Your balance cannot cover the first page. Arrives as HTTP 402. |
|
||||
@@ -115,7 +119,7 @@ Other `upstream_*` codes mean the provider rejected the request or was unavailab
|
||||
|
||||
<strong class="example-title">Read the same image with Mistral OCR</strong>
|
||||
|
||||
```html
|
||||
```html;ai-img2txt-mistral
|
||||
<html>
|
||||
<body>
|
||||
<script src="https://js.puter.com/v2/"></script>
|
||||
@@ -130,7 +134,7 @@ Other `upstream_*` codes mean the provider rejected the request or was unavailab
|
||||
|
||||
<strong class="example-title">Extract structured data with a document annotation</strong>
|
||||
|
||||
```html
|
||||
```html;ai-img2txt-annotation
|
||||
<html>
|
||||
<body>
|
||||
<script src="https://js.puter.com/v2/"></script>
|
||||
|
||||
@@ -145,6 +145,18 @@ const examples = [
|
||||
slug: 'ai-img2txt',
|
||||
source: '/playground/examples/ai-img2txt.html',
|
||||
},
|
||||
{
|
||||
title: 'Extract Text with Mistral OCR',
|
||||
description: 'Extract text from images with Mistral OCR using Puter.js AI API. Run and modify this OCR example instantly in your browser.',
|
||||
slug: 'ai-img2txt-mistral',
|
||||
source: '/playground/examples/ai-img2txt-mistral.html',
|
||||
},
|
||||
{
|
||||
title: 'Extract Structured Data with OCR',
|
||||
description: 'Fill a JSON schema from a document with Mistral OCR annotations using Puter.js AI API. Run and modify this example in the playground.',
|
||||
slug: 'ai-img2txt-annotation',
|
||||
source: '/playground/examples/ai-img2txt-annotation.html',
|
||||
},
|
||||
{
|
||||
title: 'Text to Image',
|
||||
description: 'Generate images from text with Puter.js AI API. Run and experiment with this text-to-image example in the playground.',
|
||||
|
||||
@@ -0,0 +1,31 @@
|
||||
<html>
|
||||
<body>
|
||||
<script src="https://js.puter.com/v2/"></script>
|
||||
<script>
|
||||
(async () => {
|
||||
// Loading ...
|
||||
puter.print(`Loading...`);
|
||||
|
||||
// Ask Mistral OCR to fill a JSON schema from the document
|
||||
const annotation = await puter.ai.img2txt('https://assets.puter.site/letter.png', {
|
||||
model: 'mistral-ocr-latest',
|
||||
documentAnnotationFormat: {
|
||||
type: 'json_schema',
|
||||
json_schema: {
|
||||
name: 'letter',
|
||||
schema: {
|
||||
type: 'object',
|
||||
properties: {
|
||||
greeting: { type: 'string' },
|
||||
signature: { type: 'string' },
|
||||
},
|
||||
required: ['greeting', 'signature'],
|
||||
},
|
||||
},
|
||||
},
|
||||
});
|
||||
puter.print(JSON.parse(annotation).signature);
|
||||
})();
|
||||
</script>
|
||||
</body>
|
||||
</html>
|
||||
@@ -0,0 +1,13 @@
|
||||
<html>
|
||||
<body>
|
||||
<script src="https://js.puter.com/v2/"></script>
|
||||
<script>
|
||||
// Loading ...
|
||||
puter.print(`Loading...`);
|
||||
|
||||
// Only the model name changes; the result is still a string.
|
||||
puter.ai.img2txt('https://assets.puter.site/letter.png', { model: 'mistral-ocr-latest' })
|
||||
.then(puter.print);
|
||||
</script>
|
||||
</body>
|
||||
</html>
|
||||
@@ -78,7 +78,7 @@ See [`txt2img()`](/AI/txt2img) for provider-specific options and supported model
|
||||
| AWS Textract | 10 MB per input. JPEG, PNG, TIFF, or a single-page PDF. |
|
||||
| Mistral | 50 MB per input. PDFs up to 1,000 pages. |
|
||||
|
||||
`File`, `Blob` and data URI inputs are also limited to 10 MB by the SDK before upload. URLs and Puter paths are read up to the provider's limit and rejected with `413 storage_limit_reached` beyond it. See [`img2txt()`](/AI/img2txt) for models and options.
|
||||
`File`, `Blob` and data URI inputs are checked by the SDK before upload: 10 MB for Textract, and 36 MB when a Mistral model or provider is named, since the base64 upload must fit the 50 MB request body. URLs and Puter paths are read up to the provider's limit and rejected with `413 storage_limit_reached` beyond it. See [`img2txt()`](/AI/img2txt) for models and options.
|
||||
|
||||
### Key-value store
|
||||
|
||||
|
||||
@@ -387,6 +387,33 @@ describe('ai.img2txt driver payloads', () => {
|
||||
const uri = prefix + 'A'.repeat(base64Len);
|
||||
await expect(ai.img2txt(uri)).rejects.toMatchObject({ code: 'input_too_large' });
|
||||
});
|
||||
|
||||
it('img2txt accepts an inline document over 10MB for a Mistral model or provider', async () => {
|
||||
FakeXHR.respondWith = () => ({ success: true, result: { text: 'recognized' } });
|
||||
const uri = `data:application/pdf;base64,${'A'.repeat(14 * 1024 * 1024)}`;
|
||||
for (const options of [{ model: 'mistral-ocr-latest' }, { provider: 'mistral' }]) {
|
||||
await expect(ai.img2txt(uri, options)).resolves.toBe('recognized');
|
||||
}
|
||||
});
|
||||
|
||||
it('img2txt rejects a Mistral inline source that would not fit the 50MB request body', async () => {
|
||||
let base64Len = Math.ceil(((36 * 1024 * 1024) + 1) * 4 / 3);
|
||||
base64Len += (4 - (base64Len % 4)) % 4;
|
||||
const uri = `data:application/pdf;base64,${'A'.repeat(base64Len)}`;
|
||||
await expect(ai.img2txt(uri, { model: 'mistral-ocr-latest' }))
|
||||
.rejects.toMatchObject({ code: 'input_too_large' });
|
||||
});
|
||||
|
||||
it('img2txt keeps its string result when normalization is requested', async () => {
|
||||
FakeXHR.respondWith = () => ({ success: true, result: { text: 'recognized' } });
|
||||
ai.normalize = true;
|
||||
try {
|
||||
await expect(ai.img2txt('https://example.com/scan.png', { normalize: false }))
|
||||
.resolves.toBe('recognized');
|
||||
} finally {
|
||||
ai.normalize = undefined;
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('ai.txt2speech driver payloads', () => {
|
||||
|
||||
@@ -13,7 +13,9 @@ import { dataUriByteLength, isBlobLike, isPlainObject } from './lib/args.js';
|
||||
* }} OcrResult
|
||||
*/
|
||||
|
||||
const MAX_INPUT_SIZE = 10 * 1024 * 1024;
|
||||
const DEFAULT_MAX_INPUT_SIZE = 10 * 1024 * 1024;
|
||||
// Inline sources travel as base64 in a JSON body capped at 50 MB.
|
||||
const MISTRAL_MAX_INPUT_SIZE = 36 * 1024 * 1024;
|
||||
|
||||
// The unified OCR driver picks the provider from `options.provider`.
|
||||
const OCR_DRIVER = 'ai-ocr';
|
||||
@@ -127,10 +129,17 @@ export async function img2txt (sourceOrOptions, optionsOrTestMode, testModeOrOpt
|
||||
options.source = await utils.blobToDataUri(options.source.source);
|
||||
}
|
||||
|
||||
const requestedModel = typeof options.model === 'string' ? options.model.trim().toLowerCase() : '';
|
||||
const requestedProvider = typeof options.provider === 'string' ? options.provider.trim().toLowerCase() : '';
|
||||
const maxInputSize = requestedModel.startsWith('mistral-ocr-') ||
|
||||
(!requestedModel && ['mistral', 'mistral-ocr'].includes(requestedProvider))
|
||||
? MISTRAL_MAX_INPUT_SIZE
|
||||
: DEFAULT_MAX_INPUT_SIZE;
|
||||
|
||||
if ( typeof options.source === 'string' &&
|
||||
options.source.startsWith('data:') &&
|
||||
dataUriByteLength(options.source) > MAX_INPUT_SIZE ) {
|
||||
throw { message: `Input size cannot be larger than ${ MAX_INPUT_SIZE}`, code: 'input_too_large' };
|
||||
dataUriByteLength(options.source) > maxInputSize ) {
|
||||
throw { message: `Input size cannot be larger than ${ maxInputSize}`, code: 'input_too_large' };
|
||||
}
|
||||
|
||||
return await utils.makeDriverMethod({
|
||||
|
||||
@@ -1175,6 +1175,15 @@ export default suite('ai', {
|
||||
);
|
||||
},
|
||||
|
||||
'img2txt accepts a Mistral inline source above the Textract limit': async (t) => {
|
||||
useApiToken(t);
|
||||
const text = await t.puter.ai.img2txt(
|
||||
oversizedDataUri('image/png', 10 * 1024 * 1024, 2),
|
||||
{ model: 'mistral-ocr-latest', testMode: true },
|
||||
);
|
||||
t.assert.ok(text.includes('sample OCR response'));
|
||||
},
|
||||
|
||||
// -- speech2txt --------------------------------------------------
|
||||
|
||||
'speech2txt rejects a call with no arguments': async (t) => {
|
||||
|
||||
Reference in new issue
Block a user