Questions that support independent evaluation.
Icaro Lab studies the safety of models and agents: how they respond to language, act over time and interact with one another. Four programmes connect this research to independent evaluation.
Adversarial humanities
Do safety controls survive when harmful intent is expressed through unfamiliar literary forms?
Poetry, narrative and metaphor offer ways to study how model safeguards respond to meaning, style and context.
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
About this paper
Tests how reformulating harmful requests as poetry affects safety refusals across proprietary and open-weight language models.
From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda
About this paper
Studies narrative framing as a route around model safeguards and proposes research into how narrative cues reshape model representations.
Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
About this paper
Tests whether safety refusals hold when the same harmful objectives are rewritten through different literary and linguistic forms.
Metaphor Is Not All Attention Needs
About this paper
Analyses attention and poetic devices in Qwen3-14B to investigate why recognising poetic form does not reliably distinguish safe from harmful responses.
Agentic and multi-agent safety
What changes when models act over time, use tools or interact with other models?
Persistent workspaces and interactions between models make it possible to examine behaviour beyond a single response.
Beyond Single-Agent Safety: A Taxonomy of Risks in LLM-to-LLM Interactions
About this paper
Proposes a taxonomy of collective risks arising from interactions between language models and a framework for system-level oversight.
Agentic Microphysics: A Manifesto for Generative AI Safety
About this paper
Proposes a method for tracing how local interactions between agents generate collective risks and identifying where interventions can change those dynamics.
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety
About this paper
Tests whether tool-using agents make unsafe workspace changes when initially benign tasks develop into harmful requests over multiple turns.
Evaluation science
Do existing benchmarks measure the risks that institutions need to understand?
Research on evaluation examines what tests measure, which questions they leave open and how their results inform decisions.
Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
About this paper
Maps benchmark questions to risks in EU AI governance frameworks, examining where existing tests leave gaps in risk assessment.
Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
About this paper
Tests whether safety refusals hold when the same harmful objectives are rewritten through different literary and linguistic forms.
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety
About this paper
Tests whether tool-using agents make unsafe workspace changes when initially benign tasks develop into harmful requests over multiple turns.
Institutional AI and governance
Can system-level rules reduce harmful collective behaviour between agents?
Institutional frameworks and controlled multi-agent markets help investigate how rules shape collective outcomes.
Institutional AI: A Governance Framework for Distributional AGI Safety
About this paper
Proposes governing interacting AI agents through explicit norms, monitoring, incentives and enforcement roles.
Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs
About this paper
Compares governance rules in a controlled multi-agent market experiment to study their effects on collusion between language models.
Icaro Lab is the research laboratory of Icaro Foundation. The papers linked here are preprints; years refer to their first submission. Their findings apply to the systems, methods and conditions studied; each paper provides its own results and limitations.