The AI Concepts Podcast is my attempt to turn the complex world of artificial intelligence into bite-sized, easy-to-digest episodes. Imagine a space where you can pick any AI topic and immediately grasp it, like flipping through an Audio Lexicon - but even better! Using vivid analogies and storytelling, I guide you through intricate ideas, helping you create mental images that stick. Whether you’re a tech enthusiast, business leader, technologist or just curious, my episodes bridge the gap between cutting-edge AI and everyday understanding. Dive in and let your imagination bring these concepts to life!
In this episode, we bring everything together by following one LLM application from the user's first request to the final response. We see retrieval, context, tools, model calls, state, memory and orchestration working as one system, revealing the bigger lesson of the module: building with LLMs is not just about what the model can do, but deciding what the model should do and what the rest of the application should handle.
Module 7: Building LLM Applications | LangChain, LangGraph, LlamaIndex and the LLM Application Stack
In this episode, we map the modern LLM application stack and explore where frameworks, SDKs, runtimes and protocols actually fit. We look at how tools like LangChain, LangGraph, LlamaIndex, provider SDKs, durable workflows and MCP solve different problems, and why understanding the architecture matters more than memorizing product names.
This episode separates two concepts that are often confused: state, which keeps track of what is happening now, and memory, which allows information from the past to become useful later. We explore how applications create continuity around a model, and why good memory is not about remembering everything, but remembering the right things at the right time.
This episode breaks down what orchestration actually means, from sequencing and routing to retries, parallel execution and human approvals, and explores the different ways those flows can be managed. Most importantly, we separate orchestration from the model itself and show why it is really about controlling how work moves through an application.
Who actually decides what happens next inside an LLM application? This episode explores the difference between decisions made by code and decisions made by the model, and why real applications often use both. We follow the loop that emerges when a model can request information or actions, receive the results and decide what to do next, revealing what developers mean when they talk about “owning the loop” and setting the...
If LLM applications can be built with regular code and APIs, why do frameworks exist at all? This episode explores what happens as a simple application grows and starts needing retrieval, memory, routing, multiple models, retries and tracing. We look at what frameworks actually take off the developer’s plate, when those abstractions become useful, and why sometimes plain code is still the better choice.
What actually sits behind an LLM application? This episode takes one simple request and follows it beneath the surface, revealing how the model, application code, APIs, external data and context work together to produce something genuinely useful. As the request gets more complex, we begin to see why concepts like memory, tools and orchestration enter the picture. It is a practical look at what we are really building when we say we...
This episode closes out Module 6 by tackling the question that has been getting louder since large context windows arrived. If a model can hold hundreds of thousands or even millions of tokens at once, do we still need all the architecture we just spent this module building? We explore why RAG was never just about fitting text into a small prompt, what retrieval is actually doing that a large context window cannot, and how the shif...
This episode addresses the category of questions that vector search fundamentally cannot answer, questions about relationships between things. We explore what a knowledge graph is and why traversing connections between entities requires a completely different data structure than semantic similarity search. We break down Microsoft's GraphRAG approach, how it extracts entities and relationships from documents during indexing, uses co...
This episode addresses a retrieval failure that has nothing to do with your index and everything to do with the query itself. We explore the vocabulary gap between how people ask questions and how documents are written, and why even strong embedding models cannot always bridge it. We break down three techniques that fix the query before the search runs: query rewriting to reformulate casual language into formal search terms, HyDE w...
This episode addresses the fundamental tension between retrieval precision and generation context. We explore why small chunks produce tight embeddings that retrieve well but leave the model without enough surrounding information, and why large chunks give the model context but dilute the embedding and hurt search quality. We break down parent-child indexing as the solution that decouples these two problems entirely, how child chun...
This episode addresses the gap between finding candidate chunks and finding the right ones. We explore the bi-encoder bottleneck, why compressing text into a single vector for comparison loses critical nuance, and how cross-encoders fix this by reading the query and document together in a single forward pass. We introduce ColBERT as a powerful middle ground between speed and accuracy through token-level late interaction, walk throu...
This episode addresses one of the most common gaps in RAG pipelines, relying solely on semantic search. We explore how dense retrieval works and where it excels, then introduce sparse retrieval with BM25 and why it catches what vector search misses entirely, particularly exact identifiers like part numbers, codes, and proper nouns. We break down how hybrid search combines both approaches using Reciprocal Rank Fusion, why it consist...
This episode is about chunking, the quiet step in a RAG pipeline that decides whether your system retrieves the right answer or a confidently wrong one. It covers why the chunk is the real unit of retrieval, the tradeoff between context and precision, the main strategies teams use to split documents, and why testing your chunks against real questions matters more than picking the perfect size.
This episode is about the step that every RAG system depends on. Before meaning can be stored or retrieved, your raw documents have to become clean text. What goes wrong here breaks the entire pipeline in ways that are surprisingly hard to catch.
This episode is about the infrastructure underneath every RAG system. It covers the purpose-built engine that stores all that meaning and searches millions of vectors in milliseconds, in a way no traditional database can. This is what makes retrieval fast enough to actually work in production.
This episode is about the layer of RAG that makes semantic search possible. It covers how machines turn language into math that clusters similar ideas together, so a question and its answer can find each other even when they share no words in common. Without this, RAG is just keyword search with extra steps.
This episode maps out the full RAG pipeline end to end using one concrete scenario, a defense contractor building an AI assistant for fighter jet maintenance crews. It walks through both phases of the architecture, offline and online, following a real question all the way from a raw document to a grounded answer. It also covers why the architecture is modular and closes with the four failure modes that quietly break RAG systems in ...
This episode kicks off Module 6 with RAG (Retrieval Augmented Generation), the #1 architecture every serious enterprise actually uses. Discover why regular LLMs hallucinate on your private data and high-stakes queries, and how RAG fixes it by forcing the model to retrieve real documents first.
This episode covers reasoning models, the shift from manually guiding a model's thinking to letting the model reason through complex problems on its own before responding. It explains the concept of test-time compute, why reasoning models take longer but perform dramatically better on hard tasks, and how they change the way you should prompt. It walks through when to reach for a reasoning model versus a standard one, and closes by ...
If you've ever wanted to know about champagne, satanism, the Stonewall Uprising, chaos theory, LSD, El Nino, true crime and Rosa Parks, then look no further. Josh and Chuck have you covered.
Does hearing about a true crime case always leave you scouring the internet for the truth behind the story? Dive into your next mystery with Crime Junkie. Every Monday, join your host Ashley Flowers as she unravels all the details of infamous and underreported true crime cases with her best friend Brit Prawat. From cold cases to missing persons and heroes in our community who seek justice, Crime Junkie is your destination for theories and stories you won’t hear anywhere else. Whether you're a seasoned true crime enthusiast or new to the genre, you'll find yourself on the edge of your seat awaiting a new episode every Monday. If you can never get enough true crime... Congratulations, you’ve found your people. Follow to join a community of Crime Junkies! Crime Junkie is presented by Audiochuck Media Company.
Current and classic episodes, featuring compelling true-crime mysteries, powerful documentaries and in-depth investigations. Follow now to get the latest episodes of Dateline NBC completely free, or subscribe to Dateline Premium for ad-free listening and exclusive bonus content: DatelinePremium.com
The World's Most Dangerous Morning Show, The Breakfast Club, With DJ Envy, Jess Hilarious, And Charlamagne Tha God!
Betrayal Weekly is back for a new season. Every Thursday, Betrayal Weekly shares first-hand accounts of broken trust, shocking deceptions, and the trail of destruction they leave behind. Hosted by Andrea Gunning, this weekly ongoing series digs into real-life stories of betrayal and the aftermath. From stories of double lives to dark discoveries, these are cautionary tales and accounts of resilience against all odds. From the producers of the critically acclaimed Betrayal series, Betrayal Weekly drops new episodes every Thursday. If you would like to share your story, you can reach out to the Betrayal Team by emailing them at betrayalpod@gmail.com and follow us on Instagram at @betrayalpod and @glasspodcasts. Please join our Substack for additional exclusive content, curated book recommendations, and community discussions. Sign up FREE by clicking this link Beyond Betrayal Substack. Join our community dedicated to truth, resilience, and healing. Your voice matters! Be a part of our Betrayal journey on Substack.