Token estimation for a PDF read the raw container bytes, which are mostly binary and
bear no relation to the extractable text, so a small-text PDF could be judged oversized
(or vice versa). Estimate from the extracted text instead, memoized on path+size+mtime.