Hindsight Agent Memory Passes 41,000 Stars on a Score Its Own Authors Published
Software / analysis
Hindsight Agent Memory Passes 41,000 Stars on a Score Its Own Authors Published
Vectorize's MIT-licensed memory library claims 91.4% on LongMemEval. The paper behind the number was written by the company's chief executive and a Virginia Tech professor, and the fine print matters to anyone about to adopt it.

Hindsight, an MIT-licensed memory library for AI agents from Vectorize, had 41,107 GitHub stars on Sept. 29, and the 91.4% LongMemEval score behind its pitch comes from a paper co-written by Vectorize's own chief executive.
GitHub's trending list showed the repository gaining 4,561 stars in a single day, the largest jump among the eight repositories on that list. The repository was created on Oct. 30, 2025, and lists 5,547 forks and 185 open issues.
Stars measure interest, not use. What follows is what the project actually claims, who produced the evidence, and what a team would have to accept to run it.
What Hindsight does and what it needs to run
Hindsight gives an agent three operations. Retain stores information after a language model extracts facts, entities and relationships. Recall pulls memories back through four parallel strategies: semantic vector search, BM25 keyword matching, entity and temporal graph links, and time-range filtering. Reflect runs a deeper pass that forms new connections and answers harder questions.
The README lists more than 60 integrations, among them Claude Code, Cursor, GitHub Copilot, Cline, Aider, LangGraph, LlamaIndex and CrewAI. It ships as a Docker image, a pip package, a Helm chart for Kubernetes, or a managed service called Hindsight Cloud. Pricing for the cloud service is not on the README or the project homepage.
Self-hosting has a hard requirement: PostgreSQL with pgvector, or Oracle AI Database 23ai. Intel Macs need the hindsight-all-slim variant. A team that runs neither database is adopting a database along with the library.
The project fits a pattern visible in the multi-agent harnesses that keep landing on GitHub, where the coding agent is a commodity and the state around it is the product. Nvidia, for its part, is pitching a watchdog for agents.
The 91.4% figure and the model behind it
The numbers come from "Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects," posted to arXiv on Dec. 14, 2025. The abstract reports two LongMemEval results, and they are not the same system.
With a 20B open-source model, Hindsight scores 83.6% overall, against a 39% baseline the abstract does not describe further. With what the abstract calls a larger backbone, it reaches 91.4%. On the LoCoMo benchmark it reports up to 89.61%, against 75.78% for the strongest prior open system.
- Baseline (as reported)39 %
- Hindsight, 20B open model83.6 %
- Hindsight, larger backbone91.4 %
Source: arXiv 2512.12818 abstract, Vectorize and Virginia Tech authors, accessed 2026-09-29
The headline number therefore depends on a model the abstract does not name. A team running the 20B tier should plan around 83.6%, not 91.4%.
| Benchmark | Hindsight | Comparison |
|---|---|---|
| LongMemEval, 20B open model | 83.6% | 39% baseline |
| LongMemEval, larger backbone | 91.4% | not given in abstract |
| LoCoMo, best result | 89.61% | 75.78% strongest prior open system |
VentureBeat's Dec. 16 coverage added category figures: knowledge-update questions rose from 60.3% to 84.6%, on conversations spanning up to 1.5 million tokens.
Who ran the benchmark
The README says the results were "independently reproduced by research collaborators at the Virginia Tech Sanghani Center for Artificial Intelligence and Data Analytics and The Washington Post."
The paper's author list has seven names: Chris Latimer, Nicoló Boschi, Andrew Neeser, Chris Bartholomew, Gaurav Srivastava, Xuan Wang and Naren Ramakrishnan. VentureBeat identified Latimer as co-founder and chief executive of Vectorize, and Ramakrishnan as a computer science professor at Virginia Tech and director of the Sanghani Center.
That makes the reproducing institution's director a co-author of the paper being reproduced. It does not make the result wrong. It does mean "independent" here describes collaborators, and the README's own word for them is "research collaborators." No result from a group with no stake in the outcome appeared in the sources reviewed.
The README also points to a live benchmarks page at benchmarks.hindsight.vectorize.io. Its figures are not reproduced here because that page was not read.
The claim about RAG
Latimer told VentureBeat: "RAG is on life support, and agent memory is about to kill it entirely." He also called the product "a drop-in replacement for your API calls."
Nothing in the paper abstract or the README compares Hindsight with a tuned retrieval-augmented generation pipeline on a customer's own documents. LongMemEval and LoCoMo test recall over long conversations. A support agent searching a product manual is a different task.
Ramakrishnan offered the more cautious framing in the same article: "If you have a one-size-fits-all approach to memory, either you're carrying too much context you shouldn't be carrying, or you're carrying too little context."
What is still unknown
The repository is 11 months old, and its star count did not come from a single announcement that the sources identify. Why the daily total reached 4,561 on Sept. 29 is unexplained in anything read here.
Vectorize has not published the price of Hindsight Cloud on the pages reviewed. The 185 open issues are the place to check whether the reproduction results hold on other hardware. For now, the only hard constraint in the documentation is the database: Postgres with pgvector, or Oracle 23ai.
Sources
More in Software
- 01Ponytail Hits 151,400 GitHub Stars on a Claim of 54% Less Code, Measured by Its AuthorThe plugin tells coding agents to write the minimum. Its benchmark used Claude Haiku 4.5 on one FastAPI template, four runs per ticket, and its tracker has 98 open issues.
- 02OpenDLSS-NR Reimplements Nvidia's DLSS 5 Network in Vulkan, but You Supply the WeightsThe MIT-licensed repository claims byte-for-byte parity with Nvidia's network, yet ships no weights, so the claim cannot be reproduced from the repo alone.
- 03Mozilla Shuts Down Solo AI Website Builder; All Sites Deleted Nov. 30The export ZIP leaves out image source files, Pro subscribers get prorated refunds from Oct. 1, and Mozilla points users to Wix, Squarespace, WordPress, Bolt and Lovable.
- 04IANA Says Example.com's Animated Redesign Is About Bandwidth, Not LooksKim Davies told a Google engineer the page was split to save bytes on automated traffic. Commenters measured 713 bytes of HTML plus 2.15 kB of script and are not convinced.