InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
Results and benchmarks
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models is the primary contribution described in this paper.
| Task | Dataset | Metric | Value | Source |
|---|---|---|---|---|
| Benchmarking Mitigating Over-defense Prompt Injection Guardrail | InjecGuard (w/o MOF) | hard negative accuracy. | 87.86 | paper-derived |
| Benchmarking Mitigating Over-defense Prompt Injection Guardrail | InjecGuard (w/ MOF) | hard negative accuracy. | 91.15 | paper-derived |
Audit each benchmark finding before selecting an implementation path. Evidence refs map to the disclosure below.
Evidence graph: 3 refs, 3 links.
Utility signals: depth 90/100, grounding 85/100, status high.
Implementation
Best maintained implementation now
[ACL 2025] The official implementation of the paper "PIGuard: Prompt Injection Guardrail via Mitigating Overdefense for Free".
81 stars · 9 forks · Last push Dec 4, 2025 · MIT license
- License
- CI
- Dependencies
- Docker
Official implementation from Papers with Code · Repository link is mentioned in the paper metadata · Strong overlap with paper title keywords
leolee99/injecguard is the strongest maintained implementation based on ranking signals. License is declared (MIT). Dependency/environment manifests are present.
Open leolee99/injecguard- No CI workflows detected
- Selected leolee99/injecguard as the strongest maintained implementation for new work.
- Includes dependency/environment manifest signals.
- Repository activity is within the last 24 months.
- Official repository is preserved separately as historical context.
Compare implementation paths
Compare maintenance quality, reproducibility coverage, and evidence confidence before choosing a reproduction baseline.
- Maintenance
- Stale risk
- Confidence
- High
- Reproducibility
- Moderate
- Stars
- 81
- Last push
- Dec 4, 2025 (264d)
Official implementation from Papers with Code · Repository link is mentioned in the paper metadata
- No CI pipeline detected
- No tagged releases
- No Docker setup
- Maintenance
- Stale risk
- Confidence
- High
- Reproducibility
- Moderate
- Stars
- 81
- Last push
- Dec 4, 2025 (264d)
Official implementation from Papers with Code · Repository link is mentioned in the paper metadata
- No CI pipeline detected
- No tagged releases
- No Docker setup
- Maintenance
- Recently updated
- Confidence
- Low
- Reproducibility
- Moderate
- Stars
- 12
- Last push
- Mar 24, 2026 (153d)
Matched via arXiv identifier search
- No CI pipeline detected
- No tagged releases
- No Docker setup
Reproduction readiness
Setup required
Dependencies pinned, manual setup needed
- leolee99/injecguard has requirements.txt but requires manual environment setup.
- Last push was 264 days ago, so expect possible dependency version conflicts.
- No Dockerfile, so you will set up the environment manually.
- No CI pipeline, so test coverage is unknown.
Quick start
git clone https://github.com/leolee99/injecguard.git
pip install -r requirements.txt Repositories and ecosystem
No additional verified repositories beyond the primary recommendation.
These repositories had low-confidence matching signals and are hidden by default.
- dcarpintero/pangolin-guard
Confidence: Low · 12 stars
- rakshit-737/taintwall
Confidence: Low · 0 stars
- Prateek-Pulastya/Guardrail-As-A-Service-V2
Confidence: Low · 0 stars
- tristan-kkim/guardrail4agent
Confidence: Low · 0 stars
- mrivera42/prompt-injection-classifier
Confidence: Low · 0 stars
Hugging Face artifacts
No direct paper-linked artifacts were found. Showing strongest curated related artifacts for faster exploration.
Models
No trustworthy models matches right now.
Search models on Hugging FaceDatasets
- onepaneai/faithfulness-f1score-spl-prompt-gpt-benchmarking-old
17 downloads · 0 likes · Updated May 29, 2024
- onepaneai/faithfulness-f1score-spl-prompt-benchmarking-old
16 downloads · 0 likes · Updated May 29, 2024
Broaden dataset search
Spaces
No trustworthy spaces matches right now.
Search spaces on Hugging FaceResearch context
Tasks
Benchmarking Mitigating Over-defense Prompt Injection Guardrail
Methods
None detected
Domains
None detected
Open this paper in HFEPX to review benchmark signals, evaluation modes, and human-feedback protocol context.
Open in HFEPXJump to Paper2Code search queries derived from this paper's research context.
Data includes links from Papers with Code ( CC-BY-SA-4.0 ).