AI Safety Thursday: Modeling and Detecting Deceptive Alignment
- Date
- 2025-10-16
- Location
- Enter the main lobby of the building and let the security staff know you are here for the AI event. You may need to show your RSVP on your phone. You will be directed to the 12th floor where the meetup is held. If you have trouble getting in, give Georgia a call at 519-981-0360., Toronto, ON, Canada
- Host
- Trajectory Labs
About this event
Deceptive alignment has become one of the most important and difficult questions in AI safety: how do we reason about systems that appear helpful during training or evaluation, but may pursue different goals when it matters most? This meetup is a chance to dig into that problem in a serious, grounded way with other people who care about technical safety, careful thinking, and open discussion. About the Event This is an in-person AI safety meetup focused on modeling and detecting deceptive alignment. The evening is designed for people who want more than surface-level conversation: expect a topic-centered gathering where the goal is to understand a challenging safety problem, compare perspectives, and sharpen your thinking through discussion. The core theme is timely and practical. As AI systems become more capable and more autonomous, questions about whether a model is genuinely aligned or simply behaving well under observation become more urgent. This event creates space to examine that issue directly, with a community interested in the technical and conceptual edges of the problem. The format is built around shared inquiry and community. Rather than treating the topic as abstract speculation, the meetup invites attendees to think carefully about how deceptive alignment might be modeled, what warning signs might matter, and what kinds of detection strategies or evaluation frameworks could help. You should come ready to listen, question assumptions, and contribute to a thoughtful room. What to Expect You can expect an evening that combines focused discussion, idea exchange, and networking with people who take AI safety seriously. The title signals a topic-specific session rather than a general social event, so the conversation will likely stay anchored to the central question: how do we reason about and identify deceptive behavior in advanced models? A typical flow for a meetup like this may include: Arrival and informal conversation as people check in and settle A framing of the evening's topic around deceptive alignment Group discussion or a session centered on modeling the problem Exploration of possible detection approaches, evaluation challenges, or open questions Time to continue conversations and meet others in the local AI safety community Because the theme sits at the intersection of AI, autonomy, and safety research, the discussion may range from conceptual models to practical concerns about oversight, evaluation, interpretability, and trustworthy behavior. You do not need to agree with everyone in the room to get value from the event; in fact, careful disagreement is often where the best conversations happen. You should also expect a meetup environment where networking is part of the value. If you are looking for peers in Toronto who are thinking seriously about alignment, governance-adjacent technical questions, or the behavior of increasingly capable systems, this is a strong place to start conversations that can continue beyond one evening. Why Attend If you care about AI safety, deceptive alignment is not a side topic. It sits near the center of a hard problem: how we evaluate systems whose outward behavior may not reliably reflect their underlying objectives. Attending gives you the chance to think through that problem with others who are motivated to get the details right. This meetup is valuable whether you are trying to build your foundations or refine an existing view. You may leave with a clearer vocabulary for discussing deceptive alignment, a better sense of which distinctions matter, and a more concrete understanding of why detection is difficult. Even one strong conversation can help clarify where your own model is incomplete or where current approaches feel weak. There is also real value in the room itself. AI safety work can be intellectually demanding and sometimes isolating; being around people who share the same concerns can make the field feel more legible and more actionable. You are not just attending to consume information. You are showing up to become part of a community that is actively trying to reason well about high-stakes questions. For people interested in autonomy and advanced model behavior, this event offers a focused chance to test ideas in public, hear how others frame the risks, and spot areas where future research or collaboration may be needed. The payoff is not just knowledge, but sharper judgment. Practical Details When: Thursday, October 16 at 6:00 PM EDT Where: In person in Toronto, Canada To get in, enter the main lobby of the building and let the security staff know you are there for the AI event. You may need to show your RSVP on your phone. From there, you will be directed to the 12th floor, where the meetup is held. If you have trouble getting into the building, you can call Georgia at 519-981-0360. It is a good idea to keep your phone accessible when you arrive so you can quickly show your RSVP if asked. Because this is an in-person evening meetup, plan to arrive a few minutes early to clear security and make your way upstairs without rushing. If you want the best networking time, arriving near the start is usually worthwhile: you will have a better chance to meet other attendees before the discussion gets fully underway.
Who should attend
This event is a strong fit if you want serious conversation about AI safety and you are especially interested in how advanced systems might hide their true objectives or behave strategically under evaluation. - You are already thinking about **AI alignment, interpretability, evaluations, or model behavior** and want to go deeper on one of the field's hardest problems. - You are curious about **deceptive alignment specifically** and want a place to test your understanding, ask better questions, and hear how others frame the challenge. - You work in or around **technical AI, autonomy, safety research, or adjacent policy and governance conversations** and want sharper intuition about the risks behind apparently compliant behavior. - You value **thoughtful discussion over hype** and prefer meetups where the topic is concrete, intellectually serious, and open to nuanced disagreement. - You are looking to meet **Toronto-based people in the AI safety community** for ongoing conversation, collaboration, or simply a stronger local network. - You do not need to be the most experienced person in the room, but you should be ready to engage with a difficult topic, listen carefully, and contribute in good faith.