What is Retrieval-Augmented Generation (RAG)?
Video Overview & Insights
Ready to become a certified GenAI engineer? Register now and use code IBMTechYT20 for 20% off of your exam → https://ibm.biz/BdGhCF
Wonderful explanation❤
Learn about the technology → https://ibm.biz/BdMsRT
Large language models usually give great answers, but because they're limited to the training data used to create the model. Over time they can become incomplete--or worse, generate answers that are just plain wrong. One way of improving the LLM results is called "retrieval-augmented generation" or RAG. In this video, IBM Senior Research Scientist Marina Danilevsky explains the LLM/RAG framework and how this combination delivers two big advantages, namely: the model gets the most up-to-date and trustworthy facts, and you can see where the model got its info, lending more credibility to what it generates.
Ok, but how do you guys write backwards so effortlessly?
Get weekly AI, cloud, security and sustainability industry news, events and insights. → https://ibm.biz/BdK6UY
Great explanation.
More User Perspectives
thank you for explain in simple words
@DunLiu0Watching these videos I have one question. Are these techers writing inversed lettersand sentences?
@manoflohnourAmazing explanation
@ahmadmohsin7Thank You for best Explanantion
@vijayhiremath01This was a very smart explanation. I now understand the context and why
@AngelMichael-j4jSo... Rag is currently science fiction then.... I'm looking at you Claude code
@king0vdarknessyou cant write backwards? I thought everyone could!
@aikendrum1518The video gets flipped, y'all. No writing backwards.
@blahokay1tnx
@alireza.museless. All I know now is that RAG is another entity where a LLM checks for accurate data.
So:
how the does the RAG provide the most recent data?
is RAG a LLM itself ?
Why not the LLM does that within itself ?
Not what I expected to know from this vid
Horrible teaching and explanation
@SafwanAhsan-w4gwow , very well explained
@abdelhakjebari7828wooo how she's writing in reverse??
@deepak8586Such a lucid explanation. Sticks to memory.
Thanks so much.
WOW! Perfect explanation. You have some serious skills writing backwards so we can read it on our side of the glass! Amazing job.
@lee3521Such a high quality video. Thanks for the simple yet effective explanation!
@longtaolyu1872what tool are they using for the white boarding?
@musicalbacteriaInception
@servicegrowthsystemsA lot of humans would benefit from adopting RAG like thinking in their day-to-day
@donxqxGreat video 🎉
@ANUPAGRAWALLAmazing presentation! Thank you.
@anasalbadi1792What's your kids' next question? "But how do I go from RAGs to riches?" ?
@stefanbartell1579You simply rocked❤
@meenakshisundaraar7267Incredible explanation.
@D3cr7ptExplanation was on point !!!
@ReventhSathishkumarThank you!
@florianbwdKudos to you for explaining these complex concepts so simply.
@abhinav8804Hallucination is an inherent property of how LLMs work, not a bug that can be engineered away with better data or tighter control.
Here's why:
The fundamental mechanism is the problem. LLMs generate text by predicting the most statistically probable next token given context. They are not retrieving facts from a structured database — they are pattern-matching across learned distributions. That process will always produce confident-sounding outputs that can be factually wrong.
Even perfect training data doesn't fix it. Even if a developer controls every byte of training data, the model still:
Interpolates and extrapolates between learned patterns
Has no internal mechanism to distinguish "I know this" from "I'm generating this"
Cannot verify its own outputs against ground truth at inference time
It's a probabilistic system, not a knowledge store. The weights encode statistical relationships between concepts, not discrete facts with truth values attached. A model can "know" something in one context and contradict it in another.
What can be reduced, but not eliminated:
Retrieval-Augmented Generation (RAG) grounds outputs in retrieved documents
Fine-tuning on high-quality, narrow domain data reduces hallucination in that domain
Constitutional AI and RLHF shape output behavior, but don't fix the root cause
Better architectures and scale reduce frequency, not possibility
The honest answer for any serious deployment is: hallucination must be mitigated and managed, not assumed to be solvable. Any claim that a controlled training pipeline eliminates hallucination entirely should be treated with skepticism.
TL;DR RAG = retrieving relevant text from an external source and inserting it into the LLM’s context so it can continue/answer using that information.
@Tony-dp1rlSo will the model update by itself? When the model gives a confident answer without going to external data?
@xbitroHow are you so good writing backwards?
@fluffykitten077Great tips, thanks for sharing!
@DkYadav-h5hIBM certainly knows how to make complex topics so easy and quick to understand
@muskansyed2037This was very informative, appreciate it.
@RogerMilnetif that's the case we can just google search,
isnt what this is doing in the more broader terms??
just looking up the data,
and if yes, why do we even need an LLM for this, to just convert the mend the answer according to the need?
is that it?
but do we even need this heavy of LLMs to back this up??
Great job, this video was very useful!
@TatterIsaiahFantastic explanation, thank you!
@AntoneFollett-s8fVery well explained! Great work!
@Fr333dYour videos are always so helpful, thank you!
@cvbdgdfgretreThis was super helpful, thanks!
@AlexKopeckThis was exactly what I needed, thanks!
@Shantavvamadar-b6rAwesome video, looking forward to the next one!
@WilsonFamily-qLoved this video, keep up the great work!
@ColemanTowers-d6lAwesome video, looking forward to the next one!
@DivedMalin-g3bLoved this video, keep up the great work!
@MsnsnSjdjdperfect explanation hats off
@mr_ahmad_70Awesome video, looking forward to the next one!
@EmersonSylvein