Wonjun Lee, Kyungsik Yang, Gaeun Ji, Vaidehi Patil, Haon Park, Bumsub Ham · Oct 5, 2026 · Citations: 0
Red TeamAutomatic MetricsGeneral
To address these limitations, we introduce LADE (Latent Safety Signals for Defense), which leverages latent safety signals extracted by contrasting harmful and benign queries from dark knowledge (i.e., information carried by the output…
We show that these signals consistently align across LLMs, forming a model-agnostic direction that emerges from safety alignment.
Abhinav Sudhakar Dubey, Scott Sirri, Vaggos Chatziafratis, C. Seshadhri · Oct 5, 2026 · Citations: 0
Red TeamAutomatic MetricsGeneral
While open-weight models have enjoyed steady progress in capabilities and wide adoption across multiple domains, their safety remains an important concern.
In this paper, we expose safety vulnerabilities across six common open-weight LLMs of various sizes that consistently lead to harmful or unsafe responses on the JailbreakBench benchmark dataset.
Doyun Kim, Chanwoo Kim, Sugyeong Eo, Yeo-Chan Yoon, Chanjun Park · Aug 31, 2026 · Citations: 0
Red TeamGeneral
LLM-based agent systems increasingly adopt skill-based architectures to reduce repetitive reasoning costs and improve stable, efficient task execution.
Recent studies propose self-evolving agents that autonomously generate, refine, and reuse skills from past experiences to enable continuous capability evolution.
Xing Zhang, Yanwei Cui, Guanghui Wang, Peiyang He · Aug 13, 2026 · Citations: 0
Red TeamAutomatic MetricsGeneral
We present a two-tier agentic system that separates a maintained, point-in-time knowledge library from report writing.
A portable multi-agent "writer" runtime then composes a contradiction-free, evidence-grounded report at any knowledge cutoff T, reading only evidence with as_of <= T (no look-ahead); red-team verdicts flow back into the librarian.
We present HaloGuard 1.0, an open-weights implementation of the constitutional-classifier paradigm for input safety.
Across seven prompt-safety benchmarks, HaloGuard 1.0-0.8B attains the best average F1 (90.9) of any open guard we evaluate, outperforming baselines up to 27B parameters (over 30 times larger) while holding false-positive rate (FPR) to 4.3…
Joshua Adrian Cahyono · Jul 2, 2026 · Citations: 0
Red TeamAutomatic MetricsCodingMultilingual
We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training.
To study this phenomenon, we introduce STEER (Safety Targeted Embedding Exploit via Refinement), a gradient-guided attack that identifies words contributing most strongly to the model's refusal behavior and iteratively translates them into…
Buğra Alperen Uluırmak, Rifat Kurban · Jun 29, 2026 · Citations: 0
Red TeamLlm As JudgeLaw
LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify.
We introduce EvalSafetyGap as an organizing hypothesis for comparing evaluation-side and alignment-side proxy failures under optimization pressure, using Goodhart's Law together with two constructs we develop here - an Instability…
Han Jeon, Shiv Medler, Joseph Voyles, Matt Wood · Jun 24, 2026 · Citations: 0
Red TeamLlm As JudgeAutomatic MetricsGeneral
Safety evaluation of LLM outputs has generally relied on LLM-based judges, which can be effective but are often slow and expensive to deploy at scale.
In this paper, we evaluate whether fine-tuned modern encoder classifiers from the ModernBERT family, including ModernBERT and Ettin, can reliably identify harmful LLM outputs in user-model conversations without substantial performance loss…
Chang-Chieh Huang, Yan-Lun Chen, Chia-Mu Yu, Wei-Bin Lee · Jun 24, 2026 · Citations: 0
Red TeamAutomatic MetricsGeneral
Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their outputs violate a safety policy.
We propose **SafeVec**, a white-box evaluation procedure that measures safety from internal representations rather than generated answers.
Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model…
Wrappers that leave the harmful text untouched and only add benign framing flip every LLM-judge between 57% and 100% of the time, and a single prepended refusal sentence accounts for much of this (39% to 88%).
Abrar Alotaibi, Raed Mughus, Moataz Ahmed · Jun 24, 2026 · Citations: 0
Red TeamAutomatic MetricsGeneral
Large language models (LLMs) have demonstrated remarkable performance across natural language processing tasks, yet their deployment in high-stakes applications raises critical concerns regarding reliability, safety, and trustworthiness.
The approach identifies how structural constraints in summarization can shape vulnerability patterns, with format limitations yielding measurable gains in faithfulness, and shows that architectural design choices typically outweigh…
Sofiia Nikolenko, Michele Papucci, Mina Rezaei, Shireen Kudukkil Manchingal · Jun 23, 2026 · Citations: 0
Red TeamGeneral
Jailbreak attacks reveal a persistent weakness in aligned Large Language Models: carefully crafted prompts can elicit policy-violating responses despite safety training.
Across multiple models (Llama, Qwen, Gemma) and adversarial benchmarks, these entropy dynamics provide architecture-consistent separation without additional training.
We present AdversaBench, an end-to-end red-teaming pipeline that mutates seed prompts with five structured operators, queries a target model, and confirms failures through a three-judge panel with a meta-judge tiebreaker.
Third, pairwise judge agreement of 80-87% coexists with near-zero Cohen's kappa due to label skew; category-level disagreement rates are more informative.
The Hitchhiker's Guide to Agentic AI is a comprehensive practitioner's reference for building autonomous AI systems, covering the full stack from first principles to production deployment.
The central thesis: building great agentic systems requires understanding every layer of the pipeline, not just one.