Skip to content
OpenTrain AIFor AI Companies

HFEPX · Eval paper review

Auditing Long-Term Memory Evaluation: Repeated Judging, Reader Variation, and Negative Controls

Christopher J. Chanhnourack

Published

Sep 29, 2026

Citations

0

Trust level

Low

Usefulness score

0/100 (Low)

Extraction confidence

35% (Low)

Derived from extracted protocol signals and abstract evidence.

Rater population

Not reported

Signals refreshed

Oct 2, 2026

Should you rely on this paper?

This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.

Use this as background context only. Do not make protocol decisions from this page alone.

Best use

Background context only

Use if you need

Background context only.

What to verify

Validate the exact study setup in the full paper before operational use.

Main weakness

This paper looks adjacent to evaluation work, but not like a strong protocol reference.

Human feedback signal
Not explicit
Not explicit in abstract metadata
Evaluation signal
Weak or implicit
Validate from full paper
Usefulness for eval research
0/100
Adjacent candidate

Treat as adjacent context, not a core eval-method reference.

Abstract

This report audits evaluation of a long-term-memory retrieval chain on the 500 LongMemEval-S development questions. Its strongest historical reader lane scores 479 and 475 under an adapted GPT-4o rubric; re-judging the same pass-1 answers changes three labels and yields 478. Fixed-answer knowledge-update re-scoring gives 70/72 under the upstream template and 69/72 under the modified template. Reader lanes span 93 to 479 on fixed packets; paired tests between the two strongest historical lanes establish neither superiority nor equivalence. A different-family reader, configured without client tools or operator files, scores 474, 1.0 percentage point below the headline pass (paired 95% interval [-3.0,+1.0]). Live reader request bodies were not retained. With the same requested reader label, route and judge snapshot, the full package scores 474 versus 454 for baseline sessions, a difference of +4.0 percentage points [95% interval +2.2,+6.0]. Eighteen of the 23 gains, and no losses, occur where baseline packets lacked listed evidence; this post-hoc split does not identify a component effect. In recovered LoCoMo data, token-F1 gains do not survive answer-line extraction. A negative control rejects a verifier that repairs three wrong drafts but breaks eleven correct ones. All questions were used to develop the components; no untouched holdout was evaluated. These findings do not establish a new leaderboard leader or transferable memory advantage. The A/D comparison has one pass per arm, including six reused identical-prompt outcomes, with no pinned reader snapshot; B/C and repeats remain unrun. Original headline requests cannot be reconstructed and stages 1--4 remain closed. Released artifacts support packet inspection and saved-verdict recounting and re-scoring; they do not reconstruct the method.

What we could verify

These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.

Human Feedback Types

missing

None explicit

No explicit feedback protocol extracted.

"This report audits evaluation of a long-term-memory retrieval chain on the 500 LongMemEval-S development questions."

Evaluation Modes

missing

None explicit

Validate eval design from full paper text.

"This report audits evaluation of a long-term-memory retrieval chain on the 500 LongMemEval-S development questions."

Quality Controls

partial

Adjudication

Calibration/adjudication style controls detected.

"This report audits evaluation of a long-term-memory retrieval chain on the 500 LongMemEval-S development questions."

Benchmarks / Datasets

partial

Longmemeval

Useful for quick benchmark comparison.

"This report audits evaluation of a long-term-memory retrieval chain on the 500 LongMemEval-S development questions."

Reported Metrics

missing

Not extracted

No metric anchors detected.

"This report audits evaluation of a long-term-memory retrieval chain on the 500 LongMemEval-S development questions."

Benchmarks and datasets

Longmemeval

Reported metrics

No metric terms were extracted from the available abstract.

Human feedback details
Uses human feedback
No
Feedback types
None
Rater population
Not reported
Unit of annotation
Ranking (inferred)
Expertise required
General
Evaluation details
Evaluation modes
None
Agentic eval
None
Quality controls
Adjudication
Evidence quality
Low
Use this page as
Background context only

Research brief

Metadata summary

This report audits evaluation of a long-term-memory retrieval chain on the 500 LongMemEval-S development questions.

Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.

Key takeaways

  • This report audits evaluation of a long-term-memory retrieval chain on the 500 LongMemEval-S development questions.
  • Its strongest historical reader lane scores 479 and 475 under an adapted GPT-4o rubric; re-judging the same pass-1 answers changes three labels and yields 478.
  • Fixed-answer knowledge-update re-scoring gives 70/72 under the upstream template and 69/72 under the modified template.

Researcher actions

  • Compare this paper against nearby papers in the same arXiv category before using it for protocol decisions.
  • Validate inferred eval signals (Automatic metrics) against the full paper.
  • Use related-paper links to find stronger protocol-specific references.

Caveats

  • Generated from abstract + metadata only; no PDF parsing.
  • Signals below are heuristic and may miss details reported outside the abstract.

Contribution summary

  • We evaluate an auditable long-term memory system on LongMemEval-S.
  • A grok-4.6-high reader on the same packets scores 476/474, while a maximum-reasoning-effort agentic variant regresses to 461/465.
  • A second judge agrees with the headline judge on 493/500 rows (98.6%) in each pass and scores both passes 472/500; the official judge also flips three verdicts when re-scoring byte-identical pass-1 answers.

Why it matters for eval

  • A grok-4.6-high reader on the same packets scores 476/474, while a maximum-reasoning-effort agentic variant regresses to 461/465.
  • A second judge agrees with the headline judge on 493/500 rows (98.6%) in each pass and scores both passes 472/500; the official judge also flips three verdicts when re-scoring byte-identical pass-1 answers.

Researcher checklist

  • Human feedback protocol is explicit

    No explicit human feedback protocol detected.

  • Evaluation mode is explicit

    No clear evaluation mode extracted.

  • Quality control reporting appears

    Detected: Adjudication

  • Benchmark or dataset anchors are present

    Detected: Longmemeval

  • Metric reporting is present

    No metric terms extracted.