PageIndex Hits 38,142 Stars on a Benchmark Its Maker Ran
Software / news
PageIndex Hits 38,142 Stars on a Benchmark Its Maker Ran
Vectify's MIT-licensed tree index claims 98.7 percent on FinanceBench, and its newer open-source test leaves out tables, charts and calculations.
PageIndex, a retrieval library from Vectify AI that skips vector databases in favour of a tree index an LLM reads, had 38,142 stars and 3,308 forks on GitHub as of Thursday. Its headline number, 98.7 percent accuracy on FinanceBench, comes from Vectify's own system and its own testing.
The repository is MIT-licensed, was created April 1, 2025, and listed 108 open issues when read. Instead of splitting documents into chunks and ranking them by embedding similarity, it builds a hierarchical table of contents for a PDF and has a model reason its way to the right section, "the way a human expert turns to and reads the right section of a long report", the README says.
The 98.7 percent claim
The figure traces to a Vectify blog post dated Feb. 19, 2025, about a product called Mafin 2.5, which the post says was tested on the "full benchmark dataset (100%)". The README charts it against 50 percent for vector RAG, the approach it replaces.
The post gives no method for how the vector baseline was built. It does not give the question count, and it does not say who graded the answers.
FinanceBench, by Pranab Islam, Anand Kannappan, Douwe Kiela and three co-authors, submitted Nov. 20, 2023, comprises 10,231 questions about public companies. Its authors evaluated a sample of 150 cases and reported that GPT-4-Turbo with a retrieval system "incorrectly answered or refused to answer 81% of questions". The 98.7 percent is therefore a vendor-supplied score, and nobody outside Vectify is named as having reproduced it.
What the newer open-source benchmark covers
The README's August 2026 update added a local mode to the pageindex SDK, so indexing and chat run on a user's machine with their own LLM key. Vectify published a second test for that mode, PageIndex-OSS-Benchmark: 62 lookup questions over 34 PDFs totalling 1,945 pages, drawn from MMLongBench-Doc-V2.
The README says every answer is a fact that appears in running text. Charts, tables, figures and calculations are excluded, so a wrong answer reflects a retrieval or reading failure, not a reasoning one. Vectify runs this test too.
- gpt-5.6-luna, no effort85.5 %
- gpt-5.6-luna, high96.8 %
- gpt-5.6-terra, no effort90.3 %
- gpt-5.6-terra, high100 %
- gpt-5.6-sol, medium100 %
Source: VectifyAI/PageIndex-OSS-Benchmark README (vendor-run), accessed 2026-10-01
The spread is the useful part. Moving gpt-5.6-luna from no reasoning effort to high lifts accuracy from 85.5 to 96.8 percent for $0.0031 to $0.0036 a question. The top model, gpt-5.6-sol, scores 100.0 percent at $0.0810 a question, roughly 22 times the cost of luna at high effort for 3.2 more points.
What indexing and querying cost
Building a tree runs about $0.001 a page with gpt-5.6-luna as the index model, according to the README, so a 1,000-page textbook costs a little over a dollar. Nine test PDFs from 9 to 1,098 pages indexed in roughly 13 seconds to 4.5 minutes.
On the query side, the README says passing a PDF to the model natively costs 2.1x, 3.4x, 7.8x and 16.6x more than PageIndex retrieval at 52, 85, 198 and 420 pages. At 805 pages it exceeds the model's context window. Those figures cover five PDFs and come from Vectify.
OCR and image understanding exist only in PageIndex Cloud, Vectify's hosted service, not in the open-source version. Documents that need OCR or chart reading fall outside what the open-source benchmark tests. For another open-source agent tool climbing the same list, see Nvidia's OpenShell, and for the per-token prices behind these cost figures, GPT 6.1 Sol's pricing.
Vectify has not said when an independent party will run either benchmark.
Sources
More in Software
- 01Ponytail Hits 151,400 GitHub Stars on a Claim of 54% Less Code, Measured by Its AuthorThe plugin tells coding agents to write the minimum. Its benchmark used Claude Haiku 4.5 on one FastAPI template, four runs per ticket, and its tracker has 98 open issues.
- 02OpenDLSS-NR Reimplements Nvidia's DLSS 5 Network in Vulkan, but You Supply the WeightsThe MIT-licensed repository claims byte-for-byte parity with Nvidia's network, yet ships no weights, so the claim cannot be reproduced from the repo alone.
- 03Mozilla Shuts Down Solo AI Website Builder; All Sites Deleted Nov. 30The export ZIP leaves out image source files, Pro subscribers get prorated refunds from Oct. 1, and Mozilla points users to Wix, Squarespace, WordPress, Bolt and Lovable.
- 04IANA Says Example.com's Animated Redesign Is About Bandwidth, Not LooksKim Davies told a Google engineer the page was split to save bytes on automated traffic. Commenters measured 713 bytes of HTML plus 2.15 kB of script and are not convinced.