fix(ai-ocr): check credits for the whole document before calling the provider (#4055)

* fix(ai-ocr): check credits for the whole document before calling the provider

The pre-flight checked one page's cost because the page count was only
known from the provider's answer, so a 130-page PDF passed on a balance
worth a fraction of it and drove the account negative.

The driver now estimates pages before the call: a PDF's own count (read
from its page tree, including compressed object streams), capped by a
`pages` selection; one page for images and Textract, whose synchronous
API reads single-page documents. Other documents and PDFs it can't read
count 20 pages per MB, Mistral's 1,000-page limit at its 50 MB size cap.
The estimated cost is checked and held while the provider runs, so
concurrent calls see it. Usage is still metered from the pages the
provider reports.

* fix(ai-ocr): price the credit hold at the AI cost factor

Billing scales OCR usage by the ai.cost.factor hook, but the hold was sized
at the raw cost, so a balance covering N pages passed the check and was
then charged more. Hold through reserveAiCredits like the other AI drivers.
This commit is contained in:
Daniel Salazar authored and GitHub committed 2026-10-04 00:53:19 -07:00
1 parent 3dfc70cab7
commit f533eb84ae
8 files changed
+656 -71

No files matched your search

@@ -50,6 +50,7 @@ import type { Actor } from '../../core/actor.js';
import { runWithContext } from '../../core/context.js';
import { PuterServer } from '../../server.js';
import type { MeteringService } from '../../services/metering/MeteringService.js';
import { buildPdf } from '../../testFixtures/pdf.js';
import { setupTestServer } from '../../testUtil.js';
import { generateDefaultFsentries } from '../../util/userProvisioning.js';
import type { OCRDriver } from './OCRDriver.js';
@@ -776,6 +777,185 @@ describe('OCRDriver.recognize (mistral)', () => {
});
});
// ── Credit pre-flight ───────────────────────────────────────────────
describe('OCRDriver credit pre-flight', () => {
const pageType = 'mistral-ocr:mistral-ocr-4-1:page';
const perPage = OCR_COSTS[pageType];
const pdfSource = (buffer: Buffer) => dataUrl(buffer, 'application/pdf');
const checkedCosts = () => hasCreditsSpy.mock.calls.map(([, cost]) => cost);
const outstandingHolds = (actor: Actor) =>
server.stores.creditHold.outstanding(actor.user.uuid!);
it('checks and holds credits for every page of a multi-page PDF', async () => {
const { actor } = await makeUser();
const before = await server.services.metering.getRemainingUsage(actor);
let remainingDuringCall: number | undefined;
mistralOcrProcessMock.mockImplementationOnce(async () => {
remainingDuringCall =
await server.services.metering.getRemainingUsage(actor);
return { pages: [], usageInfo: { pagesProcessed: 12 } };
});
await withActor(actor, () =>
driver.recognize({
source: pdfSource(buildPdf(12)),
provider: 'mistral',
}),
);
expect(checkedCosts()).toEqual([perPage * 12]);
// Concurrent calls see the whole document's cost while it runs.
expect(remainingDuringCall).toBe(before - perPage * 12);
expect(await outstandingHolds(actor)).toBe(0);
// Billing still follows the pages the provider reports.
const [, , count, cost] = incrementUsageSpy.mock.calls.find(
([, type]) => type === pageType,
)!;
expect(count).toBe(12);
expect(cost).toBe(perPage * 12);
});
it('prices the hold at the AI cost factor', async () => {
const { actor } = await makeUser();
const doubled = (_key: string, event: { factor: number }) => {
event.factor = 2;
};
server.clients.event.on('ai.cost.factor.*', doubled);
try {
await withActor(actor, () =>
driver.recognize({
source: pdfSource(buildPdf(3)),
provider: 'mistral',
}),
);
} finally {
server.clients.event.off('ai.cost.factor.*', doubled);
}
expect(checkedCosts()).toEqual([perPage * 3 * 2]);
});
it('refuses with 402 a PDF the balance cannot cover, even when one page fits', async () => {
const { actor } = await makeUser();
const remaining =
await server.services.metering.getRemainingUsage(actor);
const affordablePages = Math.floor(remaining / perPage);
expect(affordablePages).toBeGreaterThanOrEqual(1);
expect(affordablePages).toBeLessThan(1000);
await expect(
withActor(actor, () =>
driver.recognize({
source: pdfSource(buildPdf(affordablePages + 1)),
provider: 'mistral',
}),
),
).rejects.toMatchObject({ statusCode: 402 });
expect(mistralOcrProcessMock).not.toHaveBeenCalled();
expect(await outstandingHolds(actor)).toBe(0);
});
it('releases the hold when the provider fails', async () => {
const { actor } = await makeUser();
mistralOcrProcessMock.mockRejectedValueOnce(new Error('upstream down'));
await expect(
withActor(actor, () =>
driver.recognize({
source: pdfSource(buildPdf(3)),
provider: 'mistral',
}),
),
).rejects.toThrow('upstream down');
expect(await outstandingHolds(actor)).toBe(0);
});
it('counts PDFs packed in object streams', async () => {
const { actor } = await makeUser();
mistralOcrProcessMock.mockResolvedValueOnce({ pages: [] });
await withActor(actor, () =>
driver.recognize({
source: pdfSource(buildPdf(8, { objectStream: true })),
provider: 'mistral',
}),
);
expect(checkedCosts()).toEqual([perPage * 8]);
});
it('estimates a page selection, capped by the document', async () => {
const { actor } = await makeUser();
mistralOcrProcessMock.mockResolvedValue({ pages: [] });
const source = pdfSource(buildPdf(20));
for (const pages of [
[0, 2, 2, 5],
Array.from({ length: 50 }, (_, i) => i),
[],
]) {
await withActor(actor, () =>
driver.recognize({ source, provider: 'mistral', pages }),
);
}
// Duplicates count once; an empty selection reads the whole document.
expect(checkedCosts()).toEqual([
perPage * 3,
perPage * 20,
perPage * 20,
]);
});
it('falls back to 20 pages per MB for documents it cannot count', async () => {
const { actor } = await makeUser();
mistralOcrProcessMock.mockResolvedValue({ pages: [] });
const unreadablePdf = Buffer.alloc(1024 * 1024, 0xff);
unreadablePdf.write('%PDF-1.7\n', 'latin1');
const docx =
'application/vnd.openxmlformats-officedocument.wordprocessingml.document';
for (const source of [
pdfSource(unreadablePdf),
dataUrl(Buffer.alloc(512 * 1024), docx),
// Small unreadable input stays at one page.
pdfSource(Buffer.from('x')),
]) {
await withActor(actor, () =>
driver.recognize({ source, provider: 'mistral' }),
);
}
expect(checkedCosts()).toEqual([perPage * 20, perPage * 10, perPage]);
});
it('counts an image, and any Textract input, as one page', async () => {
const { actor } = await makeUser();
mistralOcrProcessMock.mockResolvedValue({ pages: [] });
textractSendMock.mockResolvedValue({ Blocks: [{ BlockType: 'PAGE' }] });
await withActor(actor, () =>
driver.recognize({
source: dataUrl(Buffer.alloc(2 * 1024 * 1024), 'image/png'),
provider: 'mistral',
}),
);
// Textract's synchronous API reads single-page documents only.
await withActor(actor, () =>
driver.recognize({
source: pdfSource(buildPdf(30)),
provider: 'aws-textract',
}),
);
expect(checkedCosts()).toEqual([
perPage,
OCR_COSTS['aws-textract:detect-document-text:page'],
]);
});
});
// ── Default provider selection ──────────────────────────────────────
describe('OCRDriver provider aliases', () => {
+136 -70
View File
@@ -28,6 +28,7 @@ import { Actor } from '../../core/actor.js';
import { Context } from '../../core/context.js';
import { HttpError } from '../../core/http/HttpError.js';
import { mimeFromName } from '../../util/fileSigning.js';
import type { CreditHold } from '../../services/metering/types.js';
import { PuterDriver } from '../types.js';
import {
type AiMeteringService,
@@ -40,10 +41,12 @@ import {
DEFAULT_OCR_MODEL,
findOcrModel,
OCR_MAX_INPUT_BYTES,
OCR_MAX_PAGES,
RETIRED_OCR_MODELS,
type OcrModel,
type OcrProviderId,
} from './models.js';
import { countPdfPages } from './pdfPages.js';
/**
* Driver implementing `puter-ocr` — document OCR. Two providers: •
@@ -167,6 +170,54 @@ const toMistralResponseFormat = (
};
};
const inputMimeType = (loaded: LoadedFile): string =>
loaded.mimeType ??
mimeFromName(loaded.filename) ??
'application/octet-stream';
/**
* Declared documents (PDF, DOCX, PPTX, ...) and PDF bytes are documents; images
* and untyped bytes are read as a single image.
*/
const isDocumentInput = (loaded: LoadedFile): boolean => {
const mime = inputMimeType(loaded);
return (
loaded.buffer.subarray(0, 4).toString('latin1') === '%PDF' ||
loaded.filename.toLowerCase().endsWith('.pdf') ||
(!mime.startsWith('image/') && mime !== 'application/octet-stream')
);
};
/**
* The most pages a call can bill, known before the provider runs: a PDF's page
* count, or a size-based guess for other documents, capped by a `pages`
* selection and by what the provider reads in one call.
*/
const estimateOcrPages = (
loaded: LoadedFile,
provider: OcrProviderId,
selection?: unknown,
): number => {
const maxPages = OCR_MAX_PAGES[provider];
if (maxPages <= 1 || !isDocumentInput(loaded)) return 1;
// A document we can't count is billed in proportion to its size, reaching
// the provider's page limit at its size limit (20 pages per MB for
// Mistral). Small files stay cheap; a large one can't pass as one page.
const documentPages =
countPdfPages(loaded.buffer, maxPages) ??
Math.ceil(
(loaded.buffer.length * maxPages) / OCR_MAX_INPUT_BYTES[provider],
);
const selected = Array.isArray(selection)
? new Set(
selection.filter(
(page) => Number.isInteger(page) && (page as number) >= 0,
),
).size
: 0;
return Math.max(1, Math.min(documentPages, maxPages, selected || Infinity));
};
export class OCRDriver extends PuterDriver {
readonly driverInterface = 'puter-ocr';
readonly driverName = 'ai-ocr';
@@ -337,17 +388,27 @@ export class OCRDriver extends PuterDriver {
return null;
}
async #assertCredits(actor: Actor, costPerPage: number) {
// Page count is only known once the provider answers, so pre-flight
// one page's cost and meter the real total afterward.
const hasCredits = await this.services.metering.hasEnoughCredits(
/**
* Refuse a call the balance can't cover, and hold its cost while the
* provider runs so the account's concurrent calls see it. Usage is still
* metered from the pages the provider reports.
*/
async #holdCredits(
actor: Actor,
usageType: string,
cost: number,
): Promise<CreditHold> {
// Priced at the cost factor usage is recorded at.
const hold = await this.#aiMetering.reserveAiCredits(
actor,
costPerPage,
usageType,
cost,
);
if (!hasCredits)
if (!hold)
throw new HttpError(402, 'Insufficient credits', {
legacyCode: 'insufficient_funds',
});
return hold;
}
// -- AWS Textract -------------------------------------------------
@@ -371,9 +432,6 @@ export class OCRDriver extends PuterDriver {
model: OcrModel,
actor: Actor,
) {
const costPerPage = OCR_COSTS[model.pageUsageType];
await this.#assertCredits(actor, costPerPage);
// Prefer S3 direct source if the file is FS-backed; fall back to raw bytes.
const s3Info =
loaded.fsEntry &&
@@ -401,51 +459,61 @@ export class OCRDriver extends PuterDriver {
);
};
let response;
try {
try {
response = await tryRun(Boolean(s3Info));
} catch (err) {
if (!(s3Info && err instanceof InvalidS3ObjectException))
throw err;
response = await tryRun(false);
}
} catch (err) {
if (err instanceof UnsupportedDocumentException)
throw badRequest(
'AWS Textract reads JPEG, PNG, TIFF and single-page PDF documents; use a Mistral OCR model for multi-page PDFs and other formats',
);
throw err;
}
const blocks: OcrBlock[] = [];
let pageCount = 0;
for (const block of (response.Blocks ?? []) as TextractBlock[]) {
if (block.BlockType === 'PAGE') {
pageCount += 1;
continue;
}
if (block.BlockType !== 'LINE' || !block.Text) continue;
blocks.push({
type: 'text/textract:LINE',
text: block.Text,
confidence: Number(block.Confidence ?? 0),
page: Math.max(pageCount - 1, 0),
});
}
const pages = pageCount || 1;
this.#aiMetering.incrementUsage(
const costPerPage = OCR_COSTS[model.pageUsageType];
const hold = await this.#holdCredits(
actor,
model.pageUsageType,
pages,
costPerPage * pages,
costPerPage * estimateOcrPages(loaded, model.provider),
);
return {
model: model.id,
blocks,
text: blocks.map((b) => b.text).join('\n'),
};
try {
let response;
try {
try {
response = await tryRun(Boolean(s3Info));
} catch (err) {
if (!(s3Info && err instanceof InvalidS3ObjectException))
throw err;
response = await tryRun(false);
}
} catch (err) {
if (err instanceof UnsupportedDocumentException)
throw badRequest(
'AWS Textract reads JPEG, PNG, TIFF and single-page PDF documents; use a Mistral OCR model for multi-page PDFs and other formats',
);
throw err;
}
const blocks: OcrBlock[] = [];
let pageCount = 0;
for (const block of (response.Blocks ?? []) as TextractBlock[]) {
if (block.BlockType === 'PAGE') {
pageCount += 1;
continue;
}
if (block.BlockType !== 'LINE' || !block.Text) continue;
blocks.push({
type: 'text/textract:LINE',
text: block.Text,
confidence: Number(block.Confidence ?? 0),
page: Math.max(pageCount - 1, 0),
});
}
const pages = pageCount || 1;
this.#aiMetering.incrementUsage(
actor,
model.pageUsageType,
pages,
costPerPage * pages,
);
return {
model: model.id,
blocks,
text: blocks.map((b) => b.text).join('\n'),
};
} finally {
await hold.release();
}
}
// -- Mistral OCR --------------------------------------------------
@@ -492,32 +560,30 @@ export class OCRDriver extends PuterDriver {
const annotations =
payload.documentAnnotationFormat !== undefined ||
payload.bboxAnnotationFormat !== undefined;
await this.#assertCredits(
actor,
const costPerPage =
OCR_COSTS[model.pageUsageType] +
(annotations && model.annotationUsageType
? OCR_COSTS[model.annotationUsageType]
: 0),
(annotations && model.annotationUsageType
? OCR_COSTS[model.annotationUsageType]
: 0);
const hold = await this.#holdCredits(
actor,
model.pageUsageType,
costPerPage * estimateOcrPages(loaded, model.provider, args.pages),
);
const response = await this.#mistral!.ocr.process(payload);
this.#recordMistralUsage(response, model, actor, annotations);
return this.#normalizeMistralResponse(response, model);
try {
const response = await this.#mistral!.ocr.process(payload);
this.#recordMistralUsage(response, model, actor, annotations);
return this.#normalizeMistralResponse(response, model);
} finally {
await hold.release();
}
}
#mistralBuildChunk(loaded: LoadedFile): Record<string, unknown> {
const mime =
loaded.mimeType ??
mimeFromName(loaded.filename) ??
'application/octet-stream';
// Declared documents (PDF, DOCX, PPTX, ...) and PDF bytes go as a
// document; images and untyped bytes keep the image chunk.
const isDocument =
loaded.buffer.subarray(0, 4).toString('latin1') === '%PDF' ||
loaded.filename.toLowerCase().endsWith('.pdf') ||
(!mime.startsWith('image/') && mime !== 'application/octet-stream');
const mime = inputMimeType(loaded);
const dataUrl = `data:${mime};base64,${loaded.buffer.toString('base64')}`;
if (isDocument) {
if (isDocumentInput(loaded)) {
return {
type: 'document_url',
documentUrl: dataUrl,
+6
View File
@@ -78,6 +78,12 @@ export const OCR_MAX_INPUT_BYTES: Record<OcrProviderId, number> = {
mistral: 50 * 1024 * 1024,
};
/** Most pages each provider reads in one call; Textract's sync API reads one. */
export const OCR_MAX_PAGES: Record<OcrProviderId, number> = {
'aws-textract': 1,
mistral: 1000,
};
const MODEL_BY_NAME = new Map<string, OcrModel>();
for (const model of OCR_MODELS) {
MODEL_BY_NAME.set(model.id, model);
+118
View File
@@ -0,0 +1,118 @@
/*
* Copyright (C) 2024-present Puter Technologies Inc.
*
* This file is part of Puter.
*
* Puter is free software: you can redistribute it and/or modify
* it under the terms of the GNU Affero General Public License as published
* by the Free Software Foundation, either version 3 of the License, or
* (at your option) any later version.
*
* This program is distributed in the hope that it will be useful,
* but WITHOUT ANY WARRANTY; without even the implied warranty of
* MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
* GNU Affero General Public License for more details.
*
* You should have received a copy of the GNU Affero General Public License
* along with this program. If not, see <https://www.gnu.org/licenses/>.
*/
import { describe, expect, it } from 'vitest';
import { buildPdf } from '../../testFixtures/pdf.js';
import { countPdfPages } from './pdfPages.js';
const pdf = (body: string) => Buffer.from(`%PDF-1.7\n${body}`, 'latin1');
describe('countPdfPages', () => {
it('counts the pages of a multi-page PDF', () => {
expect(countPdfPages(buildPdf(1))).toBe(1);
expect(countPdfPages(buildPdf(130))).toBe(130);
});
it('reads page objects packed in a compressed object stream', () => {
const packed = buildPdf(42, { objectStream: true });
// Nothing about the pages is visible without inflating the stream.
expect(packed.includes('/Type /Page')).toBe(false);
expect(countPdfPages(packed)).toBe(42);
});
it('reads the total from a nested page tree', () => {
expect(
countPdfPages(
pdf(
'1 0 obj\n<< /Type /Pages /Kids [2 0 R 3 0 R] /Count 7 >>\nendobj\n' +
'2 0 obj\n<< /Type /Pages /Parent 1 0 R /Kids [4 0 R] /Count 4 >>\nendobj\n' +
'3 0 obj\n<< /Type /Pages /Parent 1 0 R /Kids [5 0 R] /Count 3 >>\nendobj\n',
),
),
).toBe(7);
});
it('counts page objects when the tree understates them', () => {
const pages = Array.from(
{ length: 5 },
(_, i) => `${i + 2} 0 obj\n<</Type/Page/Parent 1 0 R>>\nendobj\n`,
).join('');
expect(
countPdfPages(
pdf(`1 0 obj\n<</Type/Pages/Count 1>>\nendobj\n${pages}`),
),
).toBe(5);
});
it('reads names written with # escapes', () => {
expect(
countPdfPages(
pdf(
'1 0 obj\n<< /Type /P#61ges /C#6funt 9 >>\nendobj\n' +
'2 0 obj\n<< /Type /P#61ge >>\nendobj\n',
),
),
).toBe(9);
});
it('lets a later definition of an object replace an earlier one', () => {
// An incremental update that shrank the tree and dropped a page.
expect(
countPdfPages(
pdf(
'1 0 obj\n<< /Type /Pages /Count 3 >>\nendobj\n' +
'2 0 obj\n<< /Type /Page >>\nendobj\n' +
'3 0 obj\n<< /Type /Page >>\nendobj\n' +
'1 0 obj\n<< /Type /Pages /Count 1 >>\nendobj\n' +
'3 0 obj\n<< /Type /Annot >>\nendobj\n',
),
),
).toBe(1);
});
it('returns null for bytes that are not a PDF', () => {
expect(countPdfPages(Buffer.from('PK\x03\x04 not a pdf'))).toBeNull();
expect(countPdfPages(Buffer.alloc(0))).toBeNull();
});
it('returns null when no pages are visible', () => {
expect(countPdfPages(pdf('garbage'))).toBeNull();
});
it('returns null when an object stream cannot be decoded', () => {
// An encrypted or corrupt stream could hide the page tree.
const unreadable = pdf(
'1 0 obj\n<< /Type /Page >>\nendobj\n' +
'2 0 obj\n<< /Type /ObjStm /N 1 /First 4 /Filter /FlateDecode >>\nstream\n' +
'not deflate data\nendstream\nendobj\n',
);
expect(countPdfPages(unreadable)).toBeNull();
const otherFilter = pdf(
'2 0 obj\n<< /Type /ObjStm /N 1 /First 4 /Filter /LZWDecode >>\nstream\n' +
'data\nendstream\nendobj\n',
);
expect(countPdfPages(otherFilter)).toBeNull();
});
it('stops at the limit', () => {
expect(countPdfPages(buildPdf(50), 10)).toBe(10);
expect(countPdfPages(buildPdf(5), 10)).toBe(5);
});
});
+146
View File
@@ -0,0 +1,146 @@
/*
* Copyright (C) 2024-present Puter Technologies Inc.
*
* This file is part of Puter.
*
* Puter is free software: you can redistribute it and/or modify
* it under the terms of the GNU Affero General Public License as published
* by the Free Software Foundation, either version 3 of the License, or
* (at your option) any later version.
*
* This program is distributed in the hope that it will be useful,
* but WITHOUT ANY WARRANTY; without even the implied warranty of
* MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
* GNU Affero General Public License for more details.
*
* You should have received a copy of the GNU Affero General Public License
* along with this program. If not, see <https://www.gnu.org/licenses/>.
*/
import { inflateSync } from 'node:zlib';
/** Total that object streams may inflate to before the count gives up. */
const MAX_INFLATED_BYTES = 64 * 1024 * 1024;
// A name ends at whitespace, a delimiter or the end of input.
const TYPE_PAGE = /\/Type[\s\0]*\/Page(?![^\s\0/<>[\]()%{}])/;
const TYPE_PAGES = /\/Type[\s\0]*\/Pages(?![^\s\0/<>[\]()%{}])/;
const TYPE_OBJECT_STREAM = /\/Type[\s\0]*\/ObjStm(?![^\s\0/<>[\]()%{}])/;
const COUNT = /\/Count[\s\0]+(\d{1,10})(?![\d.])/;
const STREAM_START = />>[\s\0]*stream(?:\r\n|\r|\n)/;
const FLATE_FILTER =
/\/Filter[\s\0]*(?:\/(?:FlateDecode|Fl)|\[[\s\0]*\/(?:FlateDecode|Fl)[\s\0]*\])/;
/** `/P#61ge` is the name `/Page`. */
const decodeNameEscapes = (dict: string): string =>
dict.includes('#')
? dict.replace(/#([0-9A-Fa-f]{2})/g, (_, hex: string) =>
String.fromCharCode(parseInt(hex, 16)),
)
: dict;
/** The objects packed in an object stream, or null when it can't be read. */
const readObjectStream = (
dict: string,
data: Buffer,
budget: number,
): { objects: Array<[number, string]>; inflated: number } | null => {
let content = data;
let inflated = 0;
if (/\/Filter/.test(dict)) {
const predictor = Number(/\/Predictor[\s\0]+(\d+)/.exec(dict)?.[1]);
if (!FLATE_FILTER.test(dict) || predictor > 1 || budget < 1)
return null;
try {
content = inflateSync(data, { maxOutputLength: budget });
} catch {
return null;
}
inflated = content.length;
}
const count = Number(/\/N[\s\0]+(\d+)/.exec(dict)?.[1]);
const first = Number(/\/First[\s\0]+(\d+)/.exec(dict)?.[1]);
const text = content.toString('latin1');
if (!Number.isInteger(count) || !(first <= text.length)) return null;
// The header is `count` pairs of object number and offset from `first`.
const header = (text.slice(0, first).match(/\d+/g) ?? []).map(Number);
if (header.length < count * 2) return null;
const objects: Array<[number, string]> = [];
for (let i = 0; i < count; i++) {
const start = first + header[i * 2 + 1]!;
const end = i + 1 < count ? first + header[i * 2 + 3]! : text.length;
objects.push([header[i * 2]!, text.slice(start, end)]);
}
return { objects, inflated };
};
/**
* Pages in a PDF, read from its page tree without a full parse: the larger of
* the tree's `/Count` and the number of page objects, since a crafted file can
* understate either. Null when the bytes show neither, including an encrypted
* or undecodable object stream that could hide them. Stops at `limit`.
*/
export function countPdfPages(pdf: Buffer, limit = Infinity): number | null {
if (!pdf.subarray(0, 1024).includes('%PDF-')) return null;
const text = pdf.toString('latin1');
// Keyed by object number, so a later definition replaces an earlier one
// as it does in an incrementally updated file.
const treeCounts = new Map<number, number>();
const pageObjects = new Set<number>();
const visit = (id: number, rawDict: string): boolean => {
const dict = decodeNameEscapes(rawDict);
treeCounts.delete(id);
pageObjects.delete(id);
if (TYPE_PAGES.test(dict)) {
const count = Number(COUNT.exec(dict)?.[1] ?? 0);
treeCounts.set(id, count);
return count >= limit;
}
if (TYPE_PAGE.test(dict)) pageObjects.add(id);
return pageObjects.size >= limit;
};
let inflated = 0;
// Bounded quantifiers keep the scan linear on crafted input.
const objectHeader =
/(?<!\d)(\d{1,10})[\s\0]{1,64}\d{1,5}[\s\0]{1,64}obj\b/g;
let header: RegExpExecArray | null;
while ((header = objectHeader.exec(text))) {
const start = objectHeader.lastIndex;
const endobj = text.indexOf('endobj', start);
const end = endobj < 0 ? text.length : endobj;
objectHeader.lastIndex = end;
const body = text.slice(start, end);
const stream = STREAM_START.exec(body);
const dict = stream ? body.slice(0, stream.index + 2) : body;
if (!TYPE_OBJECT_STREAM.test(dict)) {
if (visit(Number(header[1]), dict)) return limit;
continue;
}
const dataEnd = body.lastIndexOf('endstream');
if (!stream || dataEnd < 0) return null;
const data = pdf.subarray(
start + stream.index + stream[0].length,
start + dataEnd,
);
const packed = readObjectStream(
dict,
data,
MAX_INFLATED_BYTES - inflated,
);
if (!packed) return null;
inflated += packed.inflated;
for (const [id, objectDict] of packed.objects) {
if (visit(id, objectDict)) return limit;
}
}
let pages = pageObjects.size;
for (const count of treeCounts.values()) pages = Math.max(pages, count);
return pages > 0 ? Math.min(pages, limit) : null;
}
+67
View File
@@ -0,0 +1,67 @@
/*
* Copyright (C) 2024-present Puter Technologies Inc.
*
* This file is part of Puter.
*
* Puter is free software: you can redistribute it and/or modify
* it under the terms of the GNU Affero General Public License as published
* by the Free Software Foundation, either version 3 of the License, or
* (at your option) any later version.
*
* This program is distributed in the hope that it will be useful,
* but WITHOUT ANY WARRANTY; without even the implied warranty of
* MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
* GNU Affero General Public License for more details.
*
* You should have received a copy of the GNU Affero General Public License
* along with this program. If not, see <https://www.gnu.org/licenses/>.
*/
import { deflateSync } from 'node:zlib';
/**
* A PDF with `pageCount` blank pages: a catalog, one page tree node and the
* pages. With `objectStream`, those objects are packed into a compressed object
* stream, as most current writers do.
*/
export function buildPdf(
pageCount: number,
{ objectStream = false }: { objectStream?: boolean } = {},
): Buffer {
const kids = Array.from({ length: pageCount }, (_, i) => `${i + 3} 0 R`);
const objects = [
'<< /Type /Catalog /Pages 2 0 R >>',
`<< /Type /Pages /Kids [${kids.join(' ')}] /Count ${pageCount} >>`,
...kids.map(
() => '<< /Type /Page /Parent 2 0 R /MediaBox [0 0 612 792] >>',
),
];
const trailer = 'trailer\n<< /Root 1 0 R >>\n%%EOF\n';
if (!objectStream) {
const body = objects
.map((object, i) => `${i + 1} 0 obj\n${object}\nendobj\n`)
.join('');
return Buffer.from(`%PDF-1.4\n${body}${trailer}`, 'latin1');
}
let offset = 0;
const header: string[] = [];
for (const [i, object] of objects.entries()) {
header.push(`${i + 1} ${offset}`);
offset += object.length + 1;
}
const headerText = `${header.join(' ')}\n`;
const data = deflateSync(
Buffer.from(headerText + objects.join('\n'), 'latin1'),
);
const streamId = objects.length + 1;
return Buffer.concat([
Buffer.from(
`%PDF-1.7\n${streamId} 0 obj\n<< /Type /ObjStm /N ${objects.length} /First ${headerText.length} /Filter /FlateDecode /Length ${data.length} >>\nstream\n`,
'latin1',
),
data,
Buffer.from(`\nendstream\nendobj\n${trailer}`, 'latin1'),
]);
}
+1 -1
View File
@@ -97,7 +97,7 @@ A rejection carries the error body as the backend sent it: `{ message, code }`.
| `input_too_large` | Raised by the SDK before any request is made: a `File`, `Blob` or data URI input exceeds the selected model's limit. |
| `storage_limit_reached` | The input is larger than the model accepts (HTTP 413). |
| `bad_request` | The provider or model is unknown or retired, the model does not belong to the named provider, an option is invalid, or Textract cannot read the document. |
| `insufficient_funds` | Your balance cannot cover the first page. Arrives as HTTP 402. |
| `insufficient_funds` | Your balance cannot cover the whole document, checked before the provider runs. Arrives as HTTP 402. See [OCR limits](/rate-limits-and-quotas#ocr) for how pages are counted. |
Other `upstream_*` codes mean the provider rejected the request or was unavailable; the `message` carries the provider's reason.
+2
View File
@@ -89,6 +89,8 @@ See [`txt2img()`](/AI/txt2img) for provider-specific options and supported model
`File`, `Blob` and data URI inputs are checked by the SDK before upload: 10 MB for Textract, and 36 MB when a Mistral model or provider is named, since the base64 upload must fit the 50 MB request body. URLs and Puter paths are read up to the provider's limit and rejected with `413 storage_limit_reached` beyond it. See [`img2txt()`](/AI/img2txt) for models and options.
Before the provider runs, the balance must cover every page the call can be billed for, or it fails with `402 insufficient_funds`. A PDF counts its own pages, or only the `pages` selected when that is fewer. An image, and any Textract input, counts as one page. Other documents, and PDFs whose pages can't be read, count 20 pages per MB, up to 1,000. That amount is reserved while the call runs; the charge is for the pages actually processed.
### Key-value store
| Limit | Paid | Free | Anonymous |