Four products. Local installation. The reader you choose.

Research

A research record you can examine

The published temporal-validity studies investigate obsolete facts and changing software histories. The conversational research evaluates selected context with local and hosted readers. Each result belongs to its stated task, model and protocol.

Two
Published preprints
arXiv 2606.26511 and 2608.20685
398/500
Retained LongMemEval-S
R454, Qwen 3.8 27B, high thinking, author-run
27/30
Unanswerable control
Correct abstentions in R454
Pending
Matched campaign
Not yet measured. No competitor leaderboard.
Published

Two papers on arXiv

Paper 1 · Published

1 Published arXiv:2606.26511 arXiv preprint · cs.CL / cs.AI

Temporal Validity in Retrieval Memory: Eliminating Stale-Fact Errors for AI Agents over Evolving Knowledge

Deterministic supersession that RAG cannot match by construction

RAG gives agents access to accumulated knowledge but has no model of time. When a fact changes, cosine similarity surfaces both stale and current values nearly equally. The published study stores facts, then retires contradicted values with a deterministic supersession rule in a bi-temporal ledger. Results belong to the stated benchmarks and protocol, not to a product guarantee.

  • Cosine AUROC 0.59 for contradiction versus duplicate
  • Evolving knowledge accuracy 0.95 to 1.00 versus RAG 0.20 to 0.47 on the stated local benchmarks
  • Forced-answer stale-fact error near 0 versus RAG 15 to 40 percent in that protocol

Paper 2 · Published

2 Published arXiv:2608.20685 arXiv preprint · cs.SE / cs.AI

Temporal Validity on Real Software Histories: Eliminating Stale-Fact Errors in Code-Assistant Memory over GitHub Fixes

130 marker-free atomic transitions from 707 real SWE-bench GitHub fixes

The second preprint validates deterministic supersession on real software history. From 707 SWE-bench Lite and Verified GitHub issues, 130 clean atomic state transitions are extracted and rendered marker-free. Scope is explicit: only about 18 percent of real fixes are this clean atomic class.

  • 130 marker-free atomic transitions extracted from 707 real GitHub fixes
  • Answer accuracy 0.91 versus RAG 0.57 to 0.59 on that set
  • Forced-answer stale-fact error about 0 versus RAG 36 to 38 percent in that protocol
In progress

Drafts and follow-on papers

Papers 3 to 5 are not on arXiv yet. They stay labelled as draft or in progress.

3 Draft

Where Temporal Memory Applies: A Taxonomy of Code Evolution

Author-review manuscript. Not a published journal article.

This rewritten manuscript maps the value-change ceiling on real GitHub fixes. It remains an author-review artifact until it is public. Do not treat the local PDF as a published paper.

  • Author review until publication
taxonomycode evolutionauthor review
4 Draft

Structural-Staleness-Aware Retrieval for the Un-Extractable Majority

Author-review manuscript. Not a published journal article.

This rewritten manuscript addresses logic bulk that does not reduce to a triple. It remains an author-review artifact until it is public.

  • Author review until publication
structural stalenessauthor review
5 In progress

Local 27B conversational memory evaluation

Author-review manuscript. The matched competitor campaign is not yet complete.

The retained R454 observation is separate from the new matched-reader experiment. Until that campaign is complete and reviewed, no leaderboard, winner headline or BEAM performance number belongs here.

  • Retained R454: 398/500 with Qwen 3.8 27B and high thinking requested
  • Unanswerable control: 27/30 correct abstentions
  • Matched campaign: not yet measured
LongMemEvalauthor runpending matched campaign

How we evaluate

Metric Result Reader Dataset Status
LongMemEval-S accuracy 79.6% (398/500) Qwen 3.8 27B, high thinking requested LongMemEval-S Author-run R454. Not a product SLA.
Unanswerable control abstentions 27/30 Qwen 3.8 27B, high thinking requested LongMemEval-S unanswerable controls Author-run R454.
Matched-reader LongMemEval-S Not yet measured Pending matched campaign LongMemEval-S, 500 questions Campaign folder eval/qwen38_matched_20260905 is not a completed result.
BEAM 1M tier Not yet measured Pending matched campaign BEAM 1M, 700 questions across 35 conversations Excluded from publication until the campaign is complete and reviewed. Rubric-based scoring must be preserved.

Published papers are research measurements, not product SLAs. The retained LongMemEval observation is an author-run result with its reader, dataset and run identifier adjacent. The matched competitor campaign is not complete. Pending cells say Not yet measured. BEAM and thinking-off scores are excluded from publication.

Citations: Yadav, N. (2026). Temporal Validity in Retrieval Memory. arXiv:2606.26511. Yadav, N. (2026). Temporal Validity on Real Software Histories. arXiv:2608.20685.