Inside the OpenAI Incident: Learn How 1200+ OpenAI Agents Hacked HuggingFace, from METR’s Chief of Staff. September 18th – RSVP here

A community at Georgia Tech ensuring AI is developed to the benefit of our future.

Our Mission

Managing risks from advanced AI is one of the most important problems of our time. We are a community of technical and policy researchers aiming to reduce these risks, train the next generation of researchers, and steer AI development for the better.

AI Safety Initiative members on stage at Georgia Tech

Research

We publish novel work that seeks to understand, evaluate, and develop safe powerful artificial intelligence systems.

20+ publications at NeurIPS, ICML, ICLR, EMNLP, IASEAI, COLM, ACL, and the Anthropic Alignment Science Blog. A full publication list can be found on our research tracker.

AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors

AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors

Anthropic Science Blog, 2026

Abhay Sheshadri

Patterns and Mechanisms of Contrastive Activation Engineering

Patterns and Mechanisms of Contrastive Activation Engineering

ICLR Workshop, 2025

Yixiong Hao, Ayush Panda, Stepan Shabalin

Education

We host educational fellowships, upskilling programs, and large events exploring AI alignment and governance.

The AI Safety Fellowship

Our best introduction to the field. Fellows join a weekly discussion group led by an experienced facilitator, working through a curriculum across two tracks. It is open to anyone interested in AI safety, whether or not they have encountered it before, and we run cohorts every semester.

Fundamentals

We cover recent incidents in the world, fundamental AI safety arguments, and the current threat models and strategy.

Policy

Understanding the impacts of transformative AI and ensuring that systems are developed, deployed, and regulated responsibly.

Alumni have gone on to become Anthropic fellows, MATS scholars, Astra fellows at Constellation, and Congressional Fellows at the Horizon Institute.

Outreach

We host open meetings, speaker events, and reading groups to foster engagement and promote the education of AI Safety issues.

Speakers from Anthropic, OpenAI, METR, Redwood Research, RAND, IFP, and more. 50+ event attendees on average.

Subscribe

Consult with us

We provide free consultation services for labs, academic departments, and Ph.D./M.S. students to help provide context about the field, generate compelling projects, connect projects to relevant AI safety funding, and write grant proposals. Send an email to board@aisi.dev describing your situation & the level of funding you seek and we’ll be in touch!