Every AI system that looks good in a demo eventually has to survive contact with real data, real users, and real failure modes — and that’s where most projects stall. This session covers exactly how to close that gap.
About this event
This event is part of Data Science Dojo’s ongoing series on applied AI engineering, focused on the practical decisions behind building reliable, trustworthy AI systems in production. Where many discussions of AI stop at model capability, this one goes further into the layers that determine whether a system holds up once it’s live: data quality, grounding, evaluation, observability, and agent orchestration.
The session is hosted at the Data Science Dojo office, 5010 148th Ave NE, Suite 200A, Redmond, WA 98052, on Thursday, August 13, from 6:00 PM to 7:00 PM. Attendees will meet fellow data and AI professionals from the Seattle area over a focused evening built around one topic rather than a broad survey of trends.
Muazma Zahid is currently Group Product Manager for Google BigQuery, shaping the future of data and AI workloads in Google Cloud and helping teams build trustworthy AI systems at scale. She previously led Product Strategy and Execution for Developer and AI strategy for Azure SQL Databases and SQL Server at Microsoft, driving innovations in vector search, JSON, Data API Builder, and AI-ready database capabilities that transformed how developers build intelligent data applications. As a data and AI leader, she brings deep expertise across the full spectrum of designing and implementing large-scale, intelligent data solutions — from architecture and ingestion pipelines to analytics, ML integration, and multi-cloud optimization, with a focus on building enterprise-grade, AI-native data platforms that power the next generation of applications and agents.
Beyond product leadership, Zahid is a speaker and researcher in Biomedical Engineering with multiple international publications and awards, and a lifelong advocate for diversity and inclusion — previously serving as President of Pakistani Women in Computing (PWiC), Seattle Chapter Lead for AnitaB.org, and a leader at WomenWhoCode and the Women@Microsoft ERG.
What you will learn
- Where trust actually breaks down in trustworthy AI systems — not at the model layer, but in the data, grounding, evaluation, and orchestration layers beneath it
- How data quality issues quietly undermine even the most capable models once systems move into production
- Grounding techniques that keep model outputs accurate, relevant, and tied to real, verifiable information
- Evaluation frameworks for systematically measuring model and system performance, both before and after deployment
- Observability practices for monitoring, tracing, and debugging AI systems in production — catching drift and failures early
- Design patterns for agent orchestration — coordinating multiple models, tools, and agents reliably within larger workflows
- How these foundations fit together into a repeatable approach for building AI applications that scale without sacrificing reliability
- A framework for evaluating your own AI stack and identifying where reliability risks are most likely hiding
Why trustworthy AI systems matter
As AI systems move from demos into daily business workflows, the gap between “impressive” and “trustworthy” has become one of the most consequential engineering challenges teams face. A model that performs well in isolated testing can still produce unreliable results once it’s grounded in messy real-world data, evaluated against real user behavior, or embedded in a multi-agent pipeline. Without solid foundations in data quality, grounding, evaluation, and observability, scaling AI systems can quietly turn into scaling AI failures instead.
This is also where engineering discipline starts to matter as much as model selection. Practices like systematic evaluation, production observability, and thoughtful agent orchestration are increasingly what separate teams that can deploy trustworthy AI systems with confidence from teams constantly firefighting unexpected behavior.
For teams building their own playbooks, Data Science Dojo’s blog covers many of these themes in more depth, including data engineering practices and applied AI case studies. Google Cloud’s own BigQuery AI documentation points to similar foundations — grounding, evaluation, and reliability — as core requirements for production AI on real data platforms.
Who should attend
This event is built for data scientists, machine learning engineers, AI product managers, and engineering leads who are responsible for taking AI systems beyond the prototype stage and building trustworthy AI systems that scale. It’s a strong fit if you’re currently building or maintaining a production AI or agentic system and want practical patterns for grounding, evaluation, or orchestration rather than an introductory overview. Some familiarity with how AI models and pipelines work will help you get the most out of the discussion, but no deep technical background in any specific tool is required.