The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

Machine learning and artificial intelligence are dramatically changing the way businesses operate and people live. The TWIML AI Podcast brings the top minds and ideas from the world of ML and AI to a broad and influential community of ML/AI researchers, data scientists, engineers and tech-savvy business and IT leaders. Hosted by Sam Charrington, a sought after industry analyst, speaker, commentator and thought leader. Technologies covered include machine learning, artificial intelligence, deep learning, natural language processing, neural networks, analytics, computer science, data science and more.

Episodes

Generative Benchmarking with Kelly Hong - #728

April 23, 2025 • 54 mins

In this episode, Kelly Hong, a researcher at Chroma, joins us to discuss "Generative Benchmarking," a novel approach to evaluating retrieval systems, like RAG applications, using synthetic data. Kelly explains how traditional benchmarks like MTEB fail to represent real-world query patterns and how embedding models that perform well on public benchmarks often underperform in production. The conversation explores the two-step process...

Mark as Played

Exploring the Biology of LLMs with Circuit Tracing with Emmanuel Ameisen - #727

April 14, 2025 • 94 mins

In this episode, Emmanuel Ameisen, a research engineer at Anthropic, returns to discuss two recent papers: "Circuit Tracing: Revealing Language Model Computational Graphs" and "On the Biology of a Large Language Model." Emmanuel explains how his team developed mechanistic interpretability methods to understand the internal workings of Claude by replacing dense neural network components with sparse, interpretable alternatives. The c...

Mark as Played

Teaching LLMs to Self-Reflect with Reinforcement Learning with Maohao Shen - #726

April 8, 2025 • 51 mins

Today, we're joined by Maohao Shen, PhD student at MIT to discuss his paper, “Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search.” We dig into how Satori leverages reinforcement learning to improve language model reasoning—enabling model self-reflection, self-correction, and exploration of alternative solutions. We explore the Chain-of-Action-Thought (COAT) approach, which u...

Mark as Played

Waymo's Foundation Model for Autonomous Driving with Drago Anguelov - #725

March 31, 2025 • 69 mins

Today, we're joined by Drago Anguelov, head of AI foundations at Waymo, for a deep dive into the role of foundation models in autonomous driving. Drago shares how Waymo is leveraging large-scale machine learning, including vision-language models and generative AI techniques to improve perception, planning, and simulation for its self-driving vehicles. The conversation explores the evolution of Waymo’s research stack, their custom “...

Mark as Played

Dynamic Token Merging for Efficient Byte-level Language Models with Julie Kallini - #724

March 24, 2025 • 50 mins

Today, we're joined by Julie Kallini, PhD student at Stanford University to discuss her recent papers, “MrT5: Dynamic Token Merging for Efficient Byte-level Language Models” and “Mission: Impossible Language Models.” For the MrT5 paper, we explore the importance and failings of tokenization in large language models—including inefficient compression rates for under-resourced languages—and dig into byte-level modeling as an alternati...

Mark as Played

Scaling Up Test-Time Compute with Latent Reasoning with Jonas Geiping - #723

March 17, 2025 • 58 mins

Today, we're joined by Jonas Geiping, research group leader at Ellis Institute and the Max Planck Institute for Intelligent Systems to discuss his recent paper, “Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach.” This paper proposes a novel language model architecture which uses recurrent depth to enable “thinking in latent space.” We dig into “internal reasoning” versus “verbalized reasoning”—analogou...

Mark as Played

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought with Chengzu Li - #722

March 10, 2025 • 42 mins

Today, we're joined by Chengzu Li, PhD student at the University of Cambridge to discuss his recent paper, “Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.” We explore the motivations behind MVoT, its connection to prior work like TopViewRS, and its relation to cognitive science principles such as dual coding theory. We dig into the MVoT framework along with its various task environments—maze, mini-behavior, ...

Mark as Played

Inside s1: An o1-Style Reasoning Model That Cost Under $50 to Train with Niklas Muennighoff - #721

March 3, 2025 • 49 mins

Today, we're joined by Niklas Muennighoff, a PhD student at Stanford University, to discuss his paper, “S1: Simple Test-Time Scaling.” We explore the motivations behind S1, as well as how it compares to OpenAI's O1 and DeepSeek's R1 models. We dig into the different approaches to test-time scaling, including parallel and sequential scaling, as well as S1’s data curation process, its training recipe, and its use of model distillatio...

Mark as Played

Accelerating AI Training and Inference with AWS Trainium2 with Ron Diamant - #720

February 24, 2025 • 67 mins

Today, we're joined by Ron Diamant, chief architect for Trainium at Amazon Web Services, to discuss hardware acceleration for generative AI and the design and role of the recently released Trainium2 chip. We explore the architectural differences between Trainium and GPUs, highlighting its systolic array-based compute design, and how it balances performance across key dimensions like compute, memory bandwidth, memory capacity, and n...

Mark as Played

π0: A Foundation Model for Robotics with Sergey Levine - #719

February 18, 2025 • 52 mins

Today, we're joined by Sergey Levine, associate professor at UC Berkeley and co-founder of Physical Intelligence, to discuss π0 (pi-zero), a general-purpose robotic foundation model. We dig into the model architecture, which pairs a vision language model (VLM) with a diffusion-based action expert, and the model training "recipe," emphasizing the roles of pre-training and post-training with a diverse mixture of real-world data to en...

Mark as Played

AI Trends 2025: AI Agents and Multi-Agent Systems with Victor Dibia - #718

February 10, 2025 • 104 mins

Today we’re joined by Victor Dibia, principal research software engineer at Microsoft Research, to explore the key trends and advancements in AI agents and multi-agent systems shaping 2025 and beyond. In this episode, we discuss the unique abilities that set AI agents apart from traditional software systems–reasoning, acting, communicating, and adapting. We also examine the rise of agentic foundation models, the emergence of interf...

Mark as Played

Speculative Decoding and Efficient LLM Inference with Chris Lott - #717

February 4, 2025 • 76 mins

Today, we're joined by Chris Lott, senior director of engineering at Qualcomm AI Research to discuss accelerating large language model inference. We explore the challenges presented by the LLM encoding and decoding (aka generation) and how these interact with various hardware constraints such as FLOPS, memory footprint and memory bandwidth to limit key inference metrics such as time-to-first-token, tokens per second, and tokens per...

Mark as Played

Ensuring Privacy for Any LLM with Patricia Thaine - #716

January 28, 2025 • 51 mins

Today, we're joined by Patricia Thaine, co-founder and CEO of Private AI to discuss techniques for ensuring privacy, data minimization, and compliance when using 3rd-party large language models (LLMs) and other AI services. We explore the risks of data leakage from LLMs and embeddings, the complexities of identifying and redacting personal information across various data flows, and the approach Private AI has taken to mitigate thes...

Mark as Played

AI Engineering Pitfalls with Chip Huyen - #715

January 21, 2025 • 57 mins

Today, we're joined by Chip Huyen, independent researcher and writer to discuss her new book, “AI Engineering.” We dig into the definition of AI engineering, its key differences from traditional machine learning engineering, the common pitfalls encountered in engineering AI systems, and strategies to overcome them. We also explore how Chip defines AI agents, their current limitations and capabilities, and the critical role of effec...

Mark as Played

Evolving MLOps Platforms for Generative AI and Agents with Abhijit Bose - #714

January 13, 2025 • 58 mins

Today, we're joined by Abhijit Bose, head of enterprise AI and ML platforms at Capital One to discuss the evolution of the company’s approach and insights on Generative AI and platform best practices. In this episode, we dig into the company’s platform-centric approach to AI, and how they’ve been evolving their existing MLOps and data platforms to support the new challenges and opportunities presented by generative AI workloads and...

Mark as Played

Why Agents Are Stupid & What We Can Do About It with Dan Jeffries - #713

December 16, 2024 • 68 mins

Today, we're joined by Dan Jeffries, founder and CEO of Kentauros AI to discuss the challenges currently faced by those developing advanced AI agents. We dig into how Dan defines agents and distinguishes them from other similar uses of LLM, explore various use cases for them, and dig into ways to create smarter agentic systems. Dan shared his “big brain, little brain, tool brain” approach to tackling real-world challenges in agents...

Mark as Played

Automated Reasoning to Prevent LLM Hallucination with Byron Cook - #712

December 9, 2024 • 56 mins

Today, we're joined by Byron Cook, VP and distinguished scientist in the Automated Reasoning Group at AWS to dig into the underlying technology behind the newly announced Automated Reasoning Checks feature of Amazon Bedrock Guardrails. Automated Reasoning Checks uses mathematical proofs to help LLM users safeguard against hallucinations. We explore recent advancements in the field of automated reasoning, as well as some of the ways...

Mark as Played

AI at the Edge: Qualcomm AI Research at NeurIPS 2024 with Arash Behboodi - #711

December 3, 2024 • 54 mins

Today, we're joined by Arash Behboodi, director of engineering at Qualcomm AI Research to discuss the papers and workshops Qualcomm will be presenting at this year’s NeurIPS conference. We dig into the challenges and opportunities presented by differentiable simulation in wireless systems, the sciences, and beyond. We also explore recent work that ties conformal prediction to information theory, yielding a novel approach to incorpo...

Mark as Played

AI for Network Management with Shirley Wu - #710

November 19, 2024 • 53 mins

Today, we're joined by Shirley Wu, senior director of software engineering at Juniper Networks to discuss how machine learning and artificial intelligence are transforming network management. We explore various use cases where AI and ML are applied to enhance the quality, performance, and efficiency of networks across Juniper’s customers, including diagnosing cable degradation, proactive monitoring for coverage gaps, and real-time ...

Mark as Played

Why Your RAG System Is Broken, and How to Fix It with Jason Liu - #709

November 11, 2024 • 58 mins

Today, we're joined by Jason Liu, freelance AI consultant, advisor, and creator of the Instructor library to discuss all things retrieval-augmented generation (RAG). We dig into the tactical and strategic challenges companies face with their RAG system, the different signs Jason looks for to identify looming problems, the issues he most commonly encounters, and the steps he takes to diagnose these issues. We also cover the signific...

Mark as Played

Popular Podcasts

Stuff You Should Know

If you've ever wanted to know about champagne, satanism, the Stonewall Uprising, chaos theory, LSD, El Nino, true crime and Rosa Parks, then look no further. Josh and Chuck have you covered.

Dateline NBC

Current and classic episodes, featuring compelling true-crime mysteries, powerful documentaries and in-depth investigations. Follow now to get the latest episodes of Dateline NBC completely free, or subscribe to Dateline Premium for ad-free listening and exclusive bonus content: DatelinePremium.com

Las Culturistas with Matt Rogers and Bowen Yang

Ding dong! Join your culture consultants, Matt Rogers and Bowen Yang, on an unforgettable journey into the beating heart of CULTURE. Alongside sizzling special guests, they GET INTO the hottest pop-culture moments of the day and the formative cultural experiences that turned them into Culturistas. Produced by the Big Money Players Network and iHeartRadio.

The Bobby Bones Show

Listen to 'The Bobby Bones Show' by downloading the daily full replay.

Crime Junkie

Does hearing about a true crime case always leave you scouring the internet for the truth behind the story? Dive into your next mystery with Crime Junkie. Every Monday, join your host Ashley Flowers as she unravels all the details of infamous and underreported true crime cases with her best friend Brit Prawat. From cold cases to missing persons and heroes in our community who seek justice, Crime Junkie is your destination for theories and stories you won’t hear anywhere else. Whether you're a seasoned true crime enthusiast or new to the genre, you'll find yourself on the edge of your seat awaiting a new episode every Monday. If you can never get enough true crime... Congratulations, you’ve found your people. Follow to join a community of Crime Junkies! Crime Junkie is presented by audiochuck Media Company.

Advertise With Us

The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

Episodes

.css-14f5ked{margin:0;word-break:break-word;display:-webkit-box;-webkit-box-orient:vertical;box-orient:vertical;-webkit-line-clamp:2;overflow:hidden;}Generative Benchmarking with Kelly Hong - #728

.css-r6mb8g{margin:0;word-break:break-word;display:-webkit-box;-webkit-box-orient:vertical;box-orient:vertical;-webkit-line-clamp:1;overflow:hidden;}Exploring the Biology of LLMs with Circuit Tracing with Emmanuel Ameisen - #727

Teaching LLMs to Self-Reflect with Reinforcement Learning with Maohao Shen - #726

Waymo's Foundation Model for Autonomous Driving with Drago Anguelov - #725

Dynamic Token Merging for Efficient Byte-level Language Models with Julie Kallini - #724

Scaling Up Test-Time Compute with Latent Reasoning with Jonas Geiping - #723

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought with Chengzu Li - #722

Inside s1: An o1-Style Reasoning Model That Cost Under $50 to Train with Niklas Muennighoff - #721

Accelerating AI Training and Inference with AWS Trainium2 with Ron Diamant - #720

π0: A Foundation Model for Robotics with Sergey Levine - #719

AI Trends 2025: AI Agents and Multi-Agent Systems with Victor Dibia - #718

Speculative Decoding and Efficient LLM Inference with Chris Lott - #717

Ensuring Privacy for Any LLM with Patricia Thaine - #716

AI Engineering Pitfalls with Chip Huyen - #715

Evolving MLOps Platforms for Generative AI and Agents with Abhijit Bose - #714

Why Agents Are Stupid & What We Can Do About It with Dan Jeffries - #713

Automated Reasoning to Prevent LLM Hallucination with Byron Cook - #712

AI at the Edge: Qualcomm AI Research at NeurIPS 2024 with Arash Behboodi - #711

AI for Network Management with Shirley Wu - #710

Why Your RAG System Is Broken, and How to Fix It with Jason Liu - #709

Popular Podcasts

Generative Benchmarking with Kelly Hong - #728

Exploring the Biology of LLMs with Circuit Tracing with Emmanuel Ameisen - #727