
Berlin • November 9 & 10, 2026
Close the gap between what leadership expects and what’s actually possible at LeadDev Berlin.
As large language models move from research to production, engineering leaders are facing a new kind of challenge: how do you test, deploy, and govern systems that can behave unpredictably? At GitHub, we found that reliability wasn’t enough — trust became the core metric. Building our internal coding agents required rethinking how we validated AI output, handled uncertainty, and collaborated across disciplines.
This talk explores that journey: how we built testing pipelines for non-deterministic systems, defined “ethical success criteria,” and aligned engineers, product managers, and legal teams around shared principles for responsible AI. I’ll share the technical patterns that worked — like controlled prompt experiments, sandboxed agent evaluation, and feedback loops — as well as the cultural lessons from introducing ethical review processes into fast-moving engineering work. Attendees will leave with concrete tools for building trustworthy AI systems, a playbook for leading cross-functional conversations about safety and ethics, and practical insights for balancing innovation with accountability.
In a world where AI capabilities evolve faster than our processes, this story is a reminder that trust is built, not assumed — and that engineering leaders have a critical role in making it real.
Key takeaways:
- Learn how to test and evaluate non-deterministic LLM behavior in production systems.
- See frameworks for defining and measuring “ethical success” beyond accuracy.
- Understand how to align engineers, PMs, and legal on responsible AI principles.
- Apply lessons for balancing speed, experimentation, and accountability in AI projects.