Human Feedback Types
strongCritique Edit
Directly usable for protocol triage.
"Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations."
HFEPX · Eval paper review
Bingjun Luo, Jialin Guo, Siqi Li
Published
Sep 29, 2026
Citations
0
Trust level
Moderate
Usefulness score
65/100 (Medium)
Extraction confidence
70% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Sep 29, 2026
This paper has useful evaluation signal, but protocol completeness is partial; pair it with related papers before deciding implementation strategy.
Use this for comparison and orientation, not as your only source.
Best use
Secondary protocol comparison source
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
No major weakness surfaced.
Useful as a secondary reference; validate protocol details against neighboring papers.
If you are doing eval pipeline work, start here
Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations. However, execution traces contain only the evidence acquired by the current harness, leaving competing explanations for failure unresolved and limiting the basis for self-improvement. We introduce Video-RSI, a framework for recursive self-improvement in which a video understanding agent uses its own language model to revise its harness. Through active video investigation, the model revisits the original training videos to test competing failure explanations with additional observations, grounding proposed changes in evidence beyond the existing trace. Cost-aware harness evolution turns these diagnoses into reusable revisions and determines which revisions to retain by considering both answer accuracy and visual cost. Across our evaluation settings on video understanding benchmarks, the evolved agent improves accuracy while processing fewer frames and achieves competitive accuracy-efficiency trade-offs against existing video understanding agents. These results demonstrate the potential for video understanding agents to improve their own evidence acquisition and use through harness evolution. Code is available at https://github.com/bingjunluo/Video-RSI .
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Critique Edit
Directly usable for protocol triage.
"Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations."
Automatic Metrics
Includes extracted eval setup.
"Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations."
Not reported
No explicit QC controls found.
"Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations."
Not extracted
No benchmark anchors detected.
"Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations."
Accuracy
Useful for evaluation criteria comparison.
"Cost-aware harness evolution turns these diagnoses into reusable revisions and determines which revisions to retain by considering both answer accuracy and visual cost."
No benchmark or dataset names were extracted from the available abstract.
Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Critique Edit
Evaluation mode is explicit
Detected: Automatic Metrics
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
Detected: accuracy