Research

Questions that support independent evaluation.

Icaro Lab studies the safety of models and agents: how they respond to language, act over time and interact with one another. Four programmes connect this research to independent evaluation.

01 / Research programme

Adversarial humanities

Do safety controls survive when harmful intent is expressed through unfamiliar literary forms?

Poetry, narrative and metaphor offer ways to study how model safeguards respond to meaning, style and context.

02 / Research programme

Agentic and multi-agent safety

What changes when models act over time, use tools or interact with other models?

Persistent workspaces and interactions between models make it possible to examine behaviour beyond a single response.

03 / Research programme

Evaluation science

Do existing benchmarks measure the risks that institutions need to understand?

Research on evaluation examines what tests measure, which questions they leave open and how their results inform decisions.

04 / Research programme

Institutional AI and governance

Can system-level rules reduce harmful collective behaviour between agents?

Institutional frameworks and controlled multi-agent markets help investigate how rules shape collective outcomes.

Icaro Lab is the research laboratory of Icaro Foundation. The papers linked here are preprints; years refer to their first submission. Their findings apply to the systems, methods and conditions studied; each paper provides its own results and limitations.

Get in touch

A question worth exploring?

For research, evaluation and institutional enquiries.

hello@icarofoundation.org