No verified implementation yet

Explore-on-Graph: Incentivizing Autonomous Exploration of Large Language Models on Knowledge Graphs with Path-refined Reward Modeling

Shiqi Yan, Yubo Chen, Ruiqi Zhou, Zhengxi Yao, Shuai Chen +6 more

February 25, 2026arXiv: 2602.21728

0 repos~a few days to reproduce

Abstract

The reasoning process of Large Language Models (LLMs) is often plagued by hallucinations and missing facts in question-answering tasks. A promising solution is to ground LLMs' answers in verifiable knowledge sources, such as Knowledge Graphs (KGs). Prevailing KG-enhanced methods typically constrained LLM reasoning either by enforcing rules during generation or by imitating paths from a fixed set of demonstrations. Ho...

Results & Benchmarks

Task	Dataset	Metric	Value
Reinforcement learning	GCR	2-hop	79.0
Reinforcement learning	CWQ	Gemini-2.5-flash	67.6
Reinforcement learning	WebQSP	Gemini-2.5-pro.	78.2

Hardware Requirements

Expect multi-day setup/compute for meaningful reproduction based on current guidance.

Best Implementation

Maintained implementation evidence is not confirmed for this paper yet.

Use the Implementation Status and Reproduction Path sections below for the current action plan.

Reproduction Path

Follow this baseline workflow to decide if this paper is worth immediate prototyping.

1
Use the paper and benchmark evidence to scope a baseline reproduction plan.
2
Start from this likely method family: Reinforcement learning.
3
Track assumptions and missing details in an experiment log before coding.

Time to first repro: a few daysEstimate is based on paper-only reproduction flow

Additional Implementations

No additional verified repositories beyond the primary recommendation.

Hugging Face Artifacts

No trustworthy direct or curated related Hugging Face artifacts were found yet.

Continue with targeted Hugging Face searches:

models

arxiv:2602.21728 Explore-on-Graph Reinforcement learning

datasets

arxiv:2602.21728 Explore-on-Graph dataset Reinforcement learning benchmark

spaces

arxiv:2602.21728 Explore-on-Graph demo Reinforcement learning gradio

Research Context