Ethical and Trustworthy AI

As AI systems become increasingly integrated into everyday life, they are evolving beyond tools that merely assist users to systems that can advise, make decisions, and take autonomous actions. The emergence of foundation models and agentic AI presents new opportunities to enhance productivity, innovation, and decision-making across society. At the same time, it raises important questions about how AI should align with human values, respond to diverse societal needs, and interact responsibly with the people and communities it affects. Given the diversity of cultural, social, and individual perspectives, developing AI that can understand and operate within these contexts is a critical challenge.

Alongside these advancements, ensuring the safety, reliability, and trustworthiness of AI systems has become increasingly important. Unlike conventional software, modern AI systems are data-driven and probabilistic in nature, making their behaviour more difficult to predict, evaluate, and validate. Issues such as inaccurate outputs, vulnerability to manipulation, and performance degradation over time can limit the safe deployment of AI, particularly in high-stakes domains. Addressing these challenges requires robust methods, frameworks, and evidence-based approaches to assure AI systems before they are widely adopted.

The Ethical and Trustworthy AI (ETAI) pillar focuses on advancing trustworthy and human-centred AI that is aligned with societal values and capable of operating responsibly in complex real-world environments. By integrating expertise from artificial intelligence, cognitive and social sciences, law, and public policy, it seeks to develop AI systems that are transparent, reliable, fair, privacy-preserving, and explainable. Through fostering effective collaboration between humans and AI, this pillar aims to build a future in which AI enhances human well-being, supports informed decision-making, and contributes to a responsible Human-AI Co-Society.

Ethical and Trustworthy AI

Fig 1. Trust as a continuous cycle: values and guarantees designed in, assured by evaluation and red-teaming — deliberately attacking a system to find its failures — then deployed and monitored, with each cycle sharpening the next. Governance frames every stage rather than following them.

Ethics and trust cannot be bolted on after the fact. They must be built into how systems represent values, explain their behaviour, measure their properties, and are governed once deployed — across five mutually reinforcing thrusts:

Research Focus

We develop models that reason about values in context rather than imposing a single worldview. This spans multilingual and multicultural evaluation, elicitation and aggregation of plural preferences, and alignment techniques faithful to differing norms.

Guiding questions:

  • What values has an AI system actually learned, and how can we infer them from its representations, decisions, and multi-step trajectories? (Mechanistic Interpretability, Trajectory Analysis)
  • How do human and AI values co-evolve through repeated interaction, persuasion, and adaptation?  (Behavioural Modelling)
  • How should an AI act when values conflict or cannot be reduced to a single objective? (Machine Ethics)

We build AI whose decisions are transparent, accountable, and contestable, and that collaborates with people in ways that preserve their agency and earn warranted trust. Our interpretability research moves beyond post-hoc explanation to how models actually reason. Spoken dialogue, multimodal interaction, and controlled experiments with human participants — in the lab and in the field — ground this work in how people understand, question, and work alongside AI.

Guiding questions:

  • How can AI convey its reasoning and uncertainty so people know when to rely on it and when to override it? (Calibrated Trust)
  • How should initiative and accountability be divided so that machine capability complements rather than displaces human judgement? (Human-AI Teaming)
  • As people work with AI over time, how do we prevent erosion of the judgement that oversight depends on? (Human-Centred AI)

A claim about a model's safety or fairness is only as good as the evidence behind it. We build the evaluation methods, benchmarks, and — where the problem admits it — formal bounds that turn such claims into measurements, increasingly for physical AI acting in the open world. This extends to privacy-preserving learning, including federated and differentially private training, and to formulations that make competing requirements explicit.

Guiding questions:

  • How do we measure what a model will do when inputs are attacked or the world shifts from its training data — and how far do those measurements generalise? (Robustness Evaluation)
  • How can models learn from data that cannot be shared, without leaking what they learned from it? (Privacy-Preserving ML)
  • When accuracy, robustness, fairness, and privacy cannot be maximised together, who chooses the trade-offs, and on what basis? (Multi-Objective Learning)

We build scalable platforms to red-team, evaluate, and safeguard agentic AI. Our guardrails span the full action pipeline — inputs, memory, reasoning, tool calls, and outputs — with controls for permissions, approval, and reversibility, and detection of hallucinations before they become real-world actions.

Guiding questions:

  • How can we automatically discover failures — jailbreaks, prompt injection, goal manipulation, unsafe tool use — in systems we cannot enumerate in advance? (Red-Teaming)
  • What keeps an agent's actions within intent as it plans, remembers, and calls tools across many steps? (Runtime Guardrails)
  • How do we detect behavioural drift and emerging threats in deployed systems we did not build and cannot retrain? (Continuous Monitoring)

Principles bind only when they become requirements someone can check. Working with national and international standards, we build the operational criteria, documentation, and evidence frameworks that support procurement, audit, and post-market monitoring — extended to adaptive and agentic AI, where continuous learning makes one-off assessment insufficient.

Guiding questions:

  • How do we audit and certify AI whose training data, weights, and update history are not ours to inspect? (AI Auditing)
  • What norms and institutions are needed to govern AI agents that act and transact with each other and with people at scale? (Multi-Agent Systems)
  • As AI embeds in social institutions, how do we ensure its benefits are shared fairly and outcomes stay under human control? (Computational Social Science)
  • When autonomous research agents design experiments and report findings, who is accountable, and what does peer review become? (Scientific Integrity)

Application

We develop and validate AI solutions in high-stakes environments where reliability, safety, and trust are paramount. Through rigorous research, deployment, and evaluation, we advance technologies that address complex societal and industry challenges, translating AI research into real-world impact across these key domains:

Collaborate With Us

Advancing AI requires expertise across disciplines and a shared commitment to solving complex real-world challenges. We welcome researchers working in machine learning, multimodal and embodied AI, AI safety and assurance, computational social science, cognitive science, and human-computer interaction.

Through joint projects, postdoctoral fellowships, industry collaborations, and co-supervised graduate programmes, we bring together diverse perspectives to develop AI that is capable, trustworthy, and impactful. If these challenges inspire you, we would love to hear from you.

Get in Touch