Deep Multi-Agent Reinforcement Learning for Highway On-Ramp Merging in Mixed Traffic
Abstract
Domain fit: AI-adjacent · Paper appears method- or tooling-adjacent to AI workflows with partial ecosystem coverage.
On-ramp merging is a challenging task for autonomous vehicles (AVs), especially in mixed traffic where AVs coexist with human-driven vehicles (HDVs). In this paper, we formulate the mixed-traffic highway on-ramp merging problem as a multi-agent reinforcement learning (MARL) problem, where the AVs (on both merge lane and through lane) collaboratively learn a policy to adapt to HDVs to maximize the traffic throughput. We develop an efficient and scalable MARL framework that can be used in dynamic traffic where the communication topology could be time-varying. Parameter sharing and local rewards are exploited to foster inter-agent cooperation while achieving great scalability. An action masking scheme is employed to improve learning efficiency by filtering out invalid/unsafe actions at each step. In addition, a novel priority-based safety supervisor is developed to significantly reduce collision rate and greatly expedite the training process. A gym-like simulation environment is developed and open-sourced with three different levels of traffic densities. We exploit curriculum learning to efficiently learn harder tasks from trained models under simpler settings. Comprehensive experimental results show the proposed MARL framework consistently outperforms several state-of-the-art benchmarks.
Results and benchmarks
On-ramp merging is a challenging task for autonomous vehicles (AVs), especially in mixed traffic where AVs coexist with human-driven vehicles (HDVs).
Benchmark evidence is limited
Evidence graph: 2 refs, 1 links.
Utility signals: depth 65/100, grounding 58/100, status medium.
Implementation
No direct implementation yet
Maintained implementation evidence is not confirmed for this paper yet.
Use the implementation status and reproduction sections for the current action plan.
No verified maintained repo yet
There is no verified maintained implementation yet. Use this baseline plan to decide whether to prototype now or defer.
- No direct maintained implementation was found. Use the paper PDF and citation graph to design a baseline reproduction.
- Start from related paper: PExy: The Other Side of Exploit Kits.
- Start from this likely method family: Reinforcement learning.
Time to first repro: a few days
Recommendation evidence is currently too limited for a maintained-repo choice. Use Implementation Status and Reproduction Path for a practical baseline plan.
- Estimate is based on paper-only reproduction flow
Reproduction readiness
No repo
No verified implementation available
- No maintained repository has been identified for this paper. Check adjacent implementations or HF artifacts below.
Hardware requirements
- Expect multi-day setup/compute for meaningful reproduction based on current guidance.
Validation caveat
Hugging Face artifacts
No trustworthy direct or curated related Hugging Face artifacts were found yet. Use targeted searches to quickly locate candidate models, datasets, and demos.
Tip: start with models, then check datasets and spaces if you need evaluation data or demos.
Research context
225
Citations
85
References
Tasks
Scalability, Computer science, Supervisor, Exploit, Distributed computing, Engineering, Control and Systems Engineering, Physical Sciences
Methods
Reinforcement learning
Domains
Artificial intelligence
Related papers
- PExy: The Other Side of Exploit KitsSearch on Paper2Code
2014 · Semantic similarity
- What Is the Best Way to Work with my Supervisor?Search on Paper2Code
2023 · Semantic similarity
- CONSTRUCTION OF THE SYSTEM TO JUDGE SUPREVISOR-DOCTORAL STUDENT INTERACTIONSearch on Paper2Code
2015 · Semantic similarity
- Scalable Problem Localization for Distributed Systems: Principles and PracticesSearch on Paper2Code
2011 · Semantic similarity
- Scalable Multi-purpose Network Representation for Large Scale Distributed System SimulationSearch on Paper2Code
2012 · Semantic similarity
- Distributed and scalable message transport service for high performance multi-agent systemsSearch on Paper2Code
2004 · Semantic similarity
Open this paper in HFEPX to review benchmark signals, evaluation modes, and human-feedback protocol context.
Open in HFEPXJump to Paper2Code search queries derived from this paper's research context.