Vectorize's Hindsight Drops Its Benchmark Script for a Live Site
Software / news
Vectorize's Hindsight Drops Its Benchmark Script for a Live Site
The MIT-licensed agent memory project, which claims 91.4 percent on the LongMemEval benchmark, replaced its bundled evaluation runners with a continuously updated scoreboard in its Sept. 21 release.

Vectorize.io's open-source agent memory system, Hindsight, shipped version 0.10.1 on Sept. 21, a release that quietly dropped the project's own in-repo LoComo and LongMemEval evaluation scripts in favor of a hosted, continuously updated scoreboard at benchmarks.hindsight.vectorize.io. The change appears as one line in a changelog of more than 100 merged pull requests logged in the seven days since the prior release, v0.10.0, on Sept. 14.
Hindsight is built to give AI agents memory that improves over time rather than just recalling raw conversation history, through what its documentation calls retain, recall and reflect operations. It carries an MIT license, works with 25 or more LLM providers, and can run against an existing ChatGPT Plus, Claude Pro or GitHub Copilot subscription instead of a separate API key, according to the project's GitHub page.
The 91.4 percent figure, and who is named behind it
Vectorize, a company founded in 2024 and headquartered in Boulder, Colorado, first published Hindsight on Dec. 16, 2025, claiming a score of 91.4 percent on LongMemEval, a benchmark for long-term conversational memory. "We wanted to build an agent memory system that works like human memory," Chris Latimer, Vectorize's chief executive and co-founder, said in the announcement. The release also quoted Andrew Neeser, an applied machine learning scientist at The Washington Post, and Naren Ramakrishnan, who heads AI and machine learning for Virginia Tech's Institute for Advanced Computing. Both are listed as co-authors, alongside Latimer, on the underlying research paper posted to arXiv.
| Milestone | Date |
|---|---|
| Hindsight launches; Vectorize claims 91.4% on LongMemEval | Dec. 16, 2025 |
| Version 0.10.0 released | Sept. 14, 2026 |
| Version 0.10.1 drops in-repo benchmark scripts for a hosted dashboard | Sept. 21, 2026 |
The paper describes two techniques behind that score: TEMPR, for context-aware recall based on time and named entities, and CARA, for reflection that lets an agent learn from past successes and failures. Vectorize said its top result came from running Google's Gemini 3 Pro Preview model, with separate strong results on OpenAI's open-weight GPT-OSS 120B model. The Terminal could not independently rerun the benchmark, and Vectorize's own documentation says scores from rival agent memory vendors are self-reported, without saying whether its own figure has been checked by anyone outside the paper's author list.
A fast-moving repository
Hindsight's repository has grown to 32,201 stars and 3,572 forks since its creation in October 2025, and it gained more than 2,000 stars in a single day this week, according to GitHub's own trending count. The project lists 150 open issues. Most of the pull requests merged into v0.10.1 were narrow fixes, such as correcting how the system handles non-English text during retention and closing a database deadlock that could occur when a delete and a re-ingest ran at the same time.
One merged change folded a separate project called Hermes, described in the pull request as an agent memory plugin, directly into the Hindsight repository, expanding the scope of what a single install now covers.
What the switch to a live scoreboard signals
Replacing bundled benchmark scripts with a hosted dashboard means Hindsight's headline numbers can move without a new software release, which cuts both ways. It lets the project show fresher comparisons against competing tools, but it also means the 91.4 percent figure that made the project's reputation is no longer pinned to a specific, auditable script in the repository's own history.
Hindsight is one of several open-source projects racing to become standard infrastructure for AI agents while still working out what that infrastructure should promise. Paperclip, a rival open-source agent manager, patched a maximum-severity flaw earlier in 2026 only to leak API keys again weeks later, and OpenAI has separately found a prompt injection technique capable of copying itself between agent sessions, a risk category that applies directly to any system, like Hindsight, built to carry information from one agent session into the next.
Vectorize has not said whether the hosted benchmark site will publish the methodology behind each new score, or whether Neeser and Ramakrishnan will publish a reproduction of the result that is separate from the joint paper they co-authored with Latimer.
Sources
More in Software
- 01Vercel's ScriptC Compiles TypeScript Straight to Native CodeThe experimental compiler skips the JavaScript engine entirely and cut a simple server's cold start from Node's 61.78 milliseconds to 1.78 in an outside benchmark, though a WebKit contributor says its fallback path undercuts the whole idea.
- 02Valve Tests Pyrowave, an Open-Source GPU Streaming CodecThe codec, built by VKD3D-Proton developer Hans-Kristian Arntzen, trades up to 500 Mbit/s of bandwidth for latency low enough to rival a wired display, and Valve is testing it first on Windows, macOS and SteamOS handhelds.
- 03Madeira Runs Windows PC Games on Non-Jailbroken iPhonesDeveloper Will Faust's open-source project chains Wine, FEX-Emu and DXMT to get Cyberpunk 2077 and five other titles onto iOS, though its GPL-3.0-or-later build has to be reinstalled every seven days to dodge Apple's sideloading limit.
- 04Apple Rewrites Its Tracking Prompt for Five EU CountriesBeginning with iOS 27.2, apps in Germany, France, Italy, Poland and Romania must show a reworded App Tracking Transparency prompt, after regulators in those countries fined Apple a combined 248.6 million euros or forced a settlement over the old one.