Retrieval Augmented Generation (RAG), similar to the artistry of Remy’s Ratatouille, combines the brilliance of Large Language Models (LLMs) with the precision of information retrieval. Just as Remy layers flavors in his dish, RAG-fusion seamlessly blends vectorized documents, images, audio, and video to craft nuanced responses in AI-powered applications. And much like Anton Ego’s discerning palate, RAG-search ranking ensures that the most relevant insights rise to the top.
This session focuses on the latest architectural pattern called Retrieval Augmented Generation (RAG). With a beginner-friendly introduction to why RAG is essential. Then dive into practical implementation and design considerations. Finally, exploring different RAG variations, including RAG-fusion, multi-index, and search ranking, all of this with a pinch of "Ratatouille" (as the movie).
🔗 Conference Website: https://softwarearchitecture.live
📺 CSharp TV - Dev Streaming Destination http://csharp.tv
🌎 C# Corner - Community of Software and Data Developers
https://www.c-sharpcorner.com
#CSharpTV #CSharpCorner #CSharp #SoftwareArchitectureConf
Show More Show Less View Video Transcript
0:03
right I I am sham I work as a solution
0:05
right I I am sham I work as a solution
0:05
right I I am sham I work as a solution architect at Microsoft Netherlands um
0:09
architect at Microsoft Netherlands um
0:09
architect at Microsoft Netherlands um specifically on the app Innovation and
0:11
specifically on the app Innovation and
0:11
specifically on the app Innovation and the aisi and that's basically why this
0:15
the aisi and that's basically why this
0:15
the aisi and that's basically why this topic is pretty close to
0:17
topic is pretty close to
0:17
topic is pretty close to me um and this is our agenda today so
0:20
me um and this is our agenda today so
0:21
me um and this is our agenda today so rag uh retrieval augmented generation is
0:23
rag uh retrieval augmented generation is
0:23
rag uh retrieval augmented generation is a new kid in the block when we talk
0:25
a new kid in the block when we talk
0:25
a new kid in the block when we talk about software architecture in AI space
0:29
about software architecture in AI space
0:29
about software architecture in AI space but
0:30
but
0:30
but as easy as that it sounds there are a
0:32
as easy as that it sounds there are a
0:33
as easy as that it sounds there are a lot of nuances in this architecture
0:35
lot of nuances in this architecture
0:35
lot of nuances in this architecture pattern and what I'm going to trying to
0:37
pattern and what I'm going to trying to
0:37
pattern and what I'm going to trying to cover in the next half an hour or so is
0:39
cover in the next half an hour or so is
0:39
cover in the next half an hour or so is to go through all the nuances a little
0:42
to go through all the nuances a little
0:42
to go through all the nuances a little bit and put a resemblance with the movie
0:46
bit and put a resemblance with the movie
0:46
bit and put a resemblance with the movie rat I hope some of you or most of you
0:49
rat I hope some of you or most of you
0:49
rat I hope some of you or most of you have seen the movie where the rat is
0:51
have seen the movie where the rat is
0:51
have seen the movie where the rat is basically has a knack of cooking loves
0:53
basically has a knack of cooking loves
0:53
basically has a knack of cooking loves cooking and helps um another person to
0:57
cooking and helps um another person to
0:57
cooking and helps um another person to learn and do good in cooking
1:00
learn and do good in cooking
1:00
learn and do good in cooking um so I'm going to use all sorts of
1:02
um so I'm going to use all sorts of
1:03
um so I'm going to use all sorts of references from the movie from uh food
1:07
references from the movie from uh food
1:07
references from the movie from uh food um to explain what rag is and how
1:11
um to explain what rag is and how
1:11
um to explain what rag is and how advanced you can go with rags and the
1:13
advanced you can go with rags and the
1:13
advanced you can go with rags and the different nuances to it to be honest so
1:16
different nuances to it to be honest so
1:16
different nuances to it to be honest so we talk about food without spices we all
1:18
we talk about food without spices we all
1:18
we talk about food without spices we all know how that tastes like how do do prep
1:22
know how that tastes like how do do prep
1:22
know how that tastes like how do do prep works so chopping and assembling I'm
1:24
works so chopping and assembling I'm
1:24
works so chopping and assembling I'm going to talk about how you follow a
1:26
going to talk about how you follow a
1:26
going to talk about how you follow a recipe I'm going to talk about how you
1:28
recipe I'm going to talk about how you
1:28
recipe I'm going to talk about how you get Anton approved if you remember the
1:30
get Anton approved if you remember the
1:30
get Anton approved if you remember the movie whoever you seen it andon was a
1:33
movie whoever you seen it andon was a
1:33
movie whoever you seen it andon was a food critic who used to be very yeah
1:36
food critic who used to be very yeah
1:36
food critic who used to be very yeah strict about um writing reviews of a
1:40
strict about um writing reviews of a
1:40
strict about um writing reviews of a restaurant um then I'm going to talk
1:43
restaurant um then I'm going to talk
1:43
restaurant um then I'm going to talk about how you master your Masterpiece so
1:45
about how you master your Masterpiece so
1:45
about how you master your Masterpiece so how you fine tune it how you make sure
1:47
how you fine tune it how you make sure
1:47
how you fine tune it how you make sure your rag works the way you want it to be
1:51
your rag works the way you want it to be
1:51
your rag works the way you want it to be I'm going to talk about a few secret
1:52
I'm going to talk about a few secret
1:53
I'm going to talk about a few secret splice blend so these are basically the
1:55
splice blend so these are basically the
1:55
splice blend so these are basically the advanced rack patterns or the
1:57
advanced rack patterns or the
1:57
advanced rack patterns or the architecture patterns and lastly if we
2:00
architecture patterns and lastly if we
2:01
architecture patterns and lastly if we have time we'll talk about a few use
2:03
have time we'll talk about a few use
2:03
have time we'll talk about a few use cases um there are a lot of lot of
2:06
cases um there are a lot of lot of
2:06
cases um there are a lot of lot of content here I'm not sure how much I'm
2:08
content here I'm not sure how much I'm
2:08
content here I'm not sure how much I'm going to cover it uh put your questions
2:10
going to cover it uh put your questions
2:10
going to cover it uh put your questions in the chat if we have time later on
2:12
in the chat if we have time later on
2:12
in the chat if we have time later on we'll answer it if not I'll go back to
2:14
we'll answer it if not I'll go back to
2:14
we'll answer it if not I'll go back to the YouTube page and uh try to answer
2:16
the YouTube page and uh try to answer
2:16
the YouTube page and uh try to answer your questions as much as possible so
2:19
your questions as much as possible so
2:19
your questions as much as possible so food without proper spices so let's see
2:23
food without proper spices so let's see
2:23
food without proper spices so let's see uh we already we we are in second or
2:27
uh we already we we are in second or
2:27
uh we already we we are in second or almost the third year of uh large
2:29
almost the third year of uh large
2:29
almost the third year of uh large language model hype or flow or wave
2:33
language model hype or flow or wave
2:33
language model hype or flow or wave whatever you want to call it and we all
2:35
whatever you want to call it and we all
2:35
whatever you want to call it and we all have seen how hallucination became a
2:38
have seen how hallucination became a
2:38
have seen how hallucination became a word in our daily life so everybody who
2:41
word in our daily life so everybody who
2:41
word in our daily life so everybody who has used chpd or any sort of other AI
2:44
has used chpd or any sort of other AI
2:44
has used chpd or any sort of other AI you know Hallucination is a powerful
2:46
you know Hallucination is a powerful
2:46
you know Hallucination is a powerful thing so you need to write your commands
2:49
thing so you need to write your commands
2:49
thing so you need to write your commands or instruction in such a way that the
2:52
or instruction in such a way that the
2:52
or instruction in such a way that the llm understands and gives you back a
2:54
llm understands and gives you back a
2:54
llm understands and gives you back a proper answer so for example Good Friday
2:57
proper answer so for example Good Friday
2:57
proper answer so for example Good Friday is a holiday or not if you ask this
2:59
is a holiday or not if you ask this
2:59
is a holiday or not if you ask this question to any large language model
3:01
question to any large language model
3:01
question to any large language model it'll give you a very generic answer and
3:04
it'll give you a very generic answer and
3:04
it'll give you a very generic answer and I know in India it's a holiday but in
3:06
I know in India it's a holiday but in
3:06
I know in India it's a holiday but in Netherlands it is not a holiday but in
3:10
Netherlands it is not a holiday but in
3:10
Netherlands it is not a holiday but in Microsoft if you work as a Microsoft
3:12
Microsoft if you work as a Microsoft
3:12
Microsoft if you work as a Microsoft employee and some banks here in
3:13
employee and some banks here in
3:13
employee and some banks here in Netherlands have this day as a holiday
3:17
Netherlands have this day as a holiday
3:17
Netherlands have this day as a holiday so rest of the Netherlands not but some
3:20
so rest of the Netherlands not but some
3:20
so rest of the Netherlands not but some of us will have a holiday on Good Friday
3:23
of us will have a holiday on Good Friday
3:23
of us will have a holiday on Good Friday but if you ask the generic question to
3:25
but if you ask the generic question to
3:26
but if you ask the generic question to any large language model it will give
3:27
any large language model it will give
3:27
any large language model it will give you a very generic answer but why is
3:31
you a very generic answer but why is
3:31
you a very generic answer but why is that happens you know what's the T
3:33
that happens you know what's the T
3:33
that happens you know what's the T stands in GPT is the Transformers so the
3:38
stands in GPT is the Transformers so the
3:38
stands in GPT is the Transformers so the Transformer type of model will try to
3:42
Transformer type of model will try to
3:42
Transformer type of model will try to finish the next sentence or next bunch
3:47
finish the next sentence or next bunch
3:47
finish the next sentence or next bunch of words what you put in as an input so
3:50
of words what you put in as an input so
3:50
of words what you put in as an input so if you put in as an image you put in a
3:52
if you put in as an image you put in a
3:52
if you put in as an image you put in a text or a video it'll try to find the
3:55
text or a video it'll try to find the
3:55
text or a video it'll try to find the next possible thing and why is that
3:58
next possible thing and why is that
3:58
next possible thing and why is that because the next possible thing is what
4:00
because the next possible thing is what
4:00
because the next possible thing is what it is Stained on and mostly it is
4:02
it is Stained on and mostly it is
4:03
it is Stained on and mostly it is Stained on a lot of things but it'll try
4:04
Stained on a lot of things but it'll try
4:04
Stained on a lot of things but it'll try to find the most relevant and the high
4:07
to find the most relevant and the high
4:07
to find the most relevant and the high scored next piece of information to the
4:11
scored next piece of information to the
4:11
scored next piece of information to the input that you given to so when you get
4:14
input that you given to so when you get
4:14
input that you given to so when you get back an answer that it can be holiday in
4:16
back an answer that it can be holiday in
4:16
back an answer that it can be holiday in UK example blah blah blah it's basically
4:19
UK example blah blah blah it's basically
4:19
UK example blah blah blah it's basically trying to generalize that whole
4:21
trying to generalize that whole
4:21
trying to generalize that whole knowledge that it is trained on to be
4:23
knowledge that it is trained on to be
4:23
knowledge that it is trained on to be honest but here comes the uh the the
4:28
honest but here comes the uh the the
4:28
honest but here comes the uh the the power of rag so if you along with the
4:31
power of rag so if you along with the
4:31
power of rag so if you along with the question to the llm provide some sort of
4:34
question to the llm provide some sort of
4:34
question to the llm provide some sort of information which is very specific to
4:35
information which is very specific to
4:36
information which is very specific to your question for example if I search
4:38
your question for example if I search
4:38
your question for example if I search for all the holidays that are there in
4:41
for all the holidays that are there in
4:41
for all the holidays that are there in 2024 from our HR information because
4:45
2024 from our HR information because
4:45
2024 from our HR information because they release this information of your
4:47
they release this information of your
4:47
they release this information of your yearly holidays every company does to be
4:49
yearly holidays every company does to be
4:49
yearly holidays every company does to be honest and if you add that information
4:51
honest and if you add that information
4:51
honest and if you add that information along with your question and then send
4:54
along with your question and then send
4:54
along with your question and then send that question to any LM it will properly
4:58
that question to any LM it will properly
4:58
that question to any LM it will properly kind of reason and give you the answer
5:00
kind of reason and give you the answer
5:00
kind of reason and give you the answer that you are looking for so the only
5:03
that you are looking for so the only
5:03
that you are looking for so the only thing you change you push some sort of
5:05
thing you change you push some sort of
5:05
thing you change you push some sort of information that is very contextual for
5:07
information that is very contextual for
5:07
information that is very contextual for your case before asking the question to
5:10
your case before asking the question to
5:10
your case before asking the question to LM and it will give you a proper answer
5:12
LM and it will give you a proper answer
5:12
LM and it will give you a proper answer so in very short this is what rag is the
5:16
so in very short this is what rag is the
5:16
so in very short this is what rag is the retrieval augmented generation but as I
5:19
retrieval augmented generation but as I
5:19
retrieval augmented generation but as I said in the beginning it sounds quite
5:22
said in the beginning it sounds quite
5:22
said in the beginning it sounds quite simple and straightforward but there are
5:24
simple and straightforward but there are
5:24
simple and straightforward but there are a lot of nuances there but we'll go
5:26
a lot of nuances there but we'll go
5:26
a lot of nuances there but we'll go there as uh slowly so
5:30
there as uh slowly so
5:30
there as uh slowly so technically it sound it has two
5:32
technically it sound it has two
5:32
technically it sound it has two different workflows the top one talks
5:34
different workflows the top one talks
5:34
different workflows the top one talks about how you ingest documents and
5:37
about how you ingest documents and
5:37
about how you ingest documents and normally these ingested documents go
5:40
normally these ingested documents go
5:40
normally these ingested documents go into a vector DV but to be very honest a
5:43
into a vector DV but to be very honest a
5:43
into a vector DV but to be very honest a rag can also be based on structure data
5:46
rag can also be based on structure data
5:46
rag can also be based on structure data or a non vectorized data it's textual
5:49
or a non vectorized data it's textual
5:49
or a non vectorized data it's textual information it's all about putting the
5:51
information it's all about putting the
5:51
information it's all about putting the context
5:53
context
5:53
context before you know realizing the answer
5:55
before you know realizing the answer
5:55
before you know realizing the answer from llm so the top part talks about how
5:58
from llm so the top part talks about how
5:58
from llm so the top part talks about how you ingest data so normally there's a
6:00
you ingest data so normally there's a
6:00
you ingest data so normally there's a vector database involved you have all
6:03
vector database involved you have all
6:03
vector database involved you have all the documents you create document chunks
6:05
the documents you create document chunks
6:05
the documents you create document chunks so you make smaller pieces of documents
6:08
so you make smaller pieces of documents
6:08
so you make smaller pieces of documents then you make vectorization of those
6:11
then you make vectorization of those
6:11
then you make vectorization of those documents using an embedding model from
6:13
documents using an embedding model from
6:13
documents using an embedding model from open AI or Llama Or meta and then you
6:17
open AI or Llama Or meta and then you
6:17
open AI or Llama Or meta and then you store those embeddings in a vector
6:19
store those embeddings in a vector
6:19
store those embeddings in a vector database and Vector database is nothing
6:21
database and Vector database is nothing
6:21
database and Vector database is nothing but a normal relational data ways with
6:24
but a normal relational data ways with
6:24
but a normal relational data ways with columns and values and everything but it
6:27
columns and values and everything but it
6:27
columns and values and everything but it has a special column specific for vector
6:30
has a special column specific for vector
6:30
has a special column specific for vector emings and that's in flow 16 or 32 bits
6:35
emings and that's in flow 16 or 32 bits
6:35
emings and that's in flow 16 or 32 bits and the second part which is the lower
6:36
and the second part which is the lower
6:36
and the second part which is the lower part is basically the um retrieval uh
6:41
part is basically the um retrieval uh
6:41
part is basically the um retrieval uh workflow so if you post a question if it
6:43
workflow so if you post a question if it
6:43
workflow so if you post a question if it is a vectorized data you are searching
6:45
is a vectorized data you are searching
6:45
is a vectorized data you are searching on it will create the vectorized
6:48
on it will create the vectorized
6:48
on it will create the vectorized embedding of those questions and it will
6:51
embedding of those questions and it will
6:51
embedding of those questions and it will do a search on the vector database and
6:53
do a search on the vector database and
6:53
do a search on the vector database and once you have the relevant chunks
6:55
once you have the relevant chunks
6:55
once you have the relevant chunks retrieved from the database you will
6:58
retrieved from the database you will
6:58
retrieved from the database you will pass on to the L
7:00
pass on to the L
7:00
pass on to the L to do the reasoning along with the
7:02
to do the reasoning along with the
7:02
to do the reasoning along with the question to generate an answer so the
7:05
question to generate an answer so the
7:05
question to generate an answer so the part which retrieval is is basically
7:11
part which retrieval is is basically
7:11
part which retrieval is is basically this where you do a
7:14
this where you do a
7:14
this where you do a search the augmentation is actually this
7:20
search the augmentation is actually this
7:20
search the augmentation is actually this where you pass on the context along with
7:23
where you pass on the context along with
7:23
where you pass on the context along with the LM and generation is ultimately the
7:27
the LM and generation is ultimately the
7:27
the LM and generation is ultimately the generation of the answer
7:30
generation of the answer
7:30
generation of the answer from uh the llm using all those
7:36
context um did I skip on side no sorry
7:40
context um did I skip on side no sorry
7:40
context um did I skip on side no sorry so
7:42
so
7:42
so um now chopping and assembly what do you
7:45
um now chopping and assembly what do you
7:45
um now chopping and assembly what do you mean by chopping an assembling we all
7:47
mean by chopping an assembling we all
7:47
mean by chopping an assembling we all know if you cannot put the whole cabbage
7:49
know if you cannot put the whole cabbage
7:49
know if you cannot put the whole cabbage while you're cooking the any kind of
7:52
while you're cooking the any kind of
7:52
while you're cooking the any kind of dish you need to chop it down and you
7:54
dish you need to chop it down and you
7:54
dish you need to chop it down and you know every kind of cooking has a
7:56
know every kind of cooking has a
7:56
know every kind of cooking has a different way of chopping stuff so if
7:58
different way of chopping stuff so if
7:58
different way of chopping stuff so if you're creating using this cabbage in a
8:01
you're creating using this cabbage in a
8:01
you're creating using this cabbage in a sandwich maybe your slices will be
8:03
sandwich maybe your slices will be
8:03
sandwich maybe your slices will be bigger if you're using it in a salad
8:06
bigger if you're using it in a salad
8:06
bigger if you're using it in a salad maybe your slice will be a little bit
8:07
maybe your slice will be a little bit
8:07
maybe your slice will be a little bit smaller if you're putting it in I don't
8:09
smaller if you're putting it in I don't
8:09
smaller if you're putting it in I don't know in uh some kind of a rice dish
8:12
know in uh some kind of a rice dish
8:12
know in uh some kind of a rice dish maybe you want to chop it very thinly so
8:14
maybe you want to chop it very thinly so
8:14
maybe you want to chop it very thinly so that you can cook it pretty easily so
8:17
that you can cook it pretty easily so
8:17
that you can cook it pretty easily so chunking is related to that it depends
8:19
chunking is related to that it depends
8:19
chunking is related to that it depends on how is your search going to perform
8:23
on how is your search going to perform
8:23
on how is your search going to perform how is going to your retrieval going to
8:25
how is going to your retrieval going to
8:25
how is going to your retrieval going to work out and that's where you do need to
8:27
work out and that's where you do need to
8:28
work out and that's where you do need to chunk it and there's another reason for
8:30
chunk it and there's another reason for
8:30
chunk it and there's another reason for chunking is as well so if you have a lot
8:32
chunking is as well so if you have a lot
8:32
chunking is as well so if you have a lot of information you want to send it to an
8:34
of information you want to send it to an
8:34
of information you want to send it to an llm I know jini has a higher context
8:38
llm I know jini has a higher context
8:38
llm I know jini has a higher context range so you can send a whole
8:40
range so you can send a whole
8:40
range so you can send a whole encyclopedia to be honest but the amount
8:42
encyclopedia to be honest but the amount
8:43
encyclopedia to be honest but the amount of data you transfer over the network
8:45
of data you transfer over the network
8:45
of data you transfer over the network the amount of money you also pck and
8:47
the amount of money you also pck and
8:47
the amount of money you also pck and also it takes much time to get back the
8:51
also it takes much time to get back the
8:51
also it takes much time to get back the answer from El so you cannot send a lot
8:55
answer from El so you cannot send a lot
8:55
answer from El so you cannot send a lot of information at the same time to an
8:58
of information at the same time to an
8:58
of information at the same time to an large language morning you need to make
9:00
large language morning you need to make
9:00
large language morning you need to make it as small as possible or as effective
9:04
it as small as possible or as effective
9:04
it as small as possible or as effective as possible as well to make sure the llm
9:07
as possible as well to make sure the llm
9:07
as possible as well to make sure the llm performs
9:09
performs
9:09
performs properly and that is Loosely related to
9:13
properly and that is Loosely related to
9:13
properly and that is Loosely related to tokens so um I'm not sure how many of
9:16
tokens so um I'm not sure how many of
9:16
tokens so um I'm not sure how many of you know but U llm reasons with tokens
9:20
you know but U llm reasons with tokens
9:20
you know but U llm reasons with tokens so it accepts tokens and it generate
9:22
so it accepts tokens and it generate
9:22
so it accepts tokens and it generate tokens and you can Loosely say token is
9:26
tokens and you can Loosely say token is
9:26
tokens and you can Loosely say token is a a word or a a traye or even um a bunch
9:33
a a word or a a traye or even um a bunch
9:33
a a word or a a traye or even um a bunch a bunch of characters also and uh this
9:37
a bunch of characters also and uh this
9:37
a bunch of characters also and uh this token has every llm has a limit of token
9:40
token has every llm has a limit of token
9:40
token has every llm has a limit of token that it can uh take it as an input and
9:43
that it can uh take it as an input and
9:43
that it can uh take it as an input and provide as an output so this total token
9:45
provide as an output so this total token
9:45
provide as an output so this total token limit also comes into the picture when
9:48
limit also comes into the picture when
9:48
limit also comes into the picture when you basically try to chunk your
9:51
you basically try to chunk your
9:51
you basically try to chunk your documents to ingest so remember I'm
9:54
documents to ingest so remember I'm
9:54
documents to ingest so remember I'm still talking about your ingestion part
9:56
still talking about your ingestion part
9:56
still talking about your ingestion part so the way you making your ingestion
9:58
so the way you making your ingestion
9:59
so the way you making your ingestion part
10:00
part
10:00
part um smarter or efficient the way your
10:03
um smarter or efficient the way your
10:03
um smarter or efficient the way your retrieval part will also be efficient um
10:06
retrieval part will also be efficient um
10:07
retrieval part will also be efficient um I also have an example here let me see
10:08
I also have an example here let me see
10:08
I also have an example here let me see if I can open
10:13
this so this is a free application where
10:16
this so this is a free application where
10:16
this so this is a free application where you can actually play around how your um
10:20
you can actually play around how your um
10:20
you can actually play around how your um uh how your um chunking looks like so
10:24
uh how your um chunking looks like so
10:24
uh how your um chunking looks like so normally chunking can also be talked as
10:26
normally chunking can also be talked as
10:26
normally chunking can also be talked as a splitter and there is a different kind
10:28
a splitter and there is a different kind
10:28
a splitter and there is a different kind of splitters available here this is just
10:30
of splitters available here this is just
10:31
of splitters available here this is just a small list there are multiple others
10:34
a small list there are multiple others
10:34
a small list there are multiple others just to give you an
10:37
example if I make my chunk size around
10:40
example if I make my chunk size around
10:40
example if I make my chunk size around 675 so this is my first chunk this is
10:44
675 so this is my first chunk this is
10:44
675 so this is my first chunk this is green is in my second chunk and the red
10:47
green is in my second chunk and the red
10:47
green is in my second chunk and the red and so on so forth there's also
10:49
and so on so forth there's also
10:49
and so on so forth there's also something called an overlap because when
10:51
something called an overlap because when
10:51
something called an overlap because when you are reasoning of distinctive data
10:55
you are reasoning of distinctive data
10:55
you are reasoning of distinctive data which were connected before so for
10:57
which were connected before so for
10:57
which were connected before so for example this was a whole paragraph now
10:59
example this was a whole paragraph now
10:59
example this was a whole paragraph now I'm breaking it into different chunks
11:02
I'm breaking it into different chunks
11:02
I'm breaking it into different chunks you also want to have some sort of
11:05
you also want to have some sort of
11:05
you also want to have some sort of overlap what would the overlap gives you
11:08
overlap what would the overlap gives you
11:08
overlap what would the overlap gives you the overlap gives you the piece of
11:10
the overlap gives you the piece of
11:10
the overlap gives you the piece of information that might want to flow to
11:14
information that might want to flow to
11:14
information that might want to flow to the next chunk of your uh information
11:17
the next chunk of your uh information
11:17
the next chunk of your uh information and this helps when you want to um have
11:21
and this helps when you want to um have
11:21
and this helps when you want to um have multiple search results back and you
11:23
multiple search results back and you
11:23
multiple search results back and you want to combine those chunks somehow to
11:26
want to combine those chunks somehow to
11:26
want to combine those chunks somehow to generate a reasonable information then
11:29
generate a reasonable information then
11:29
generate a reasonable information then makes sense to have these kind of
11:31
makes sense to have these kind of
11:31
makes sense to have these kind of overlap and this is like the old radios
11:33
overlap and this is like the old radios
11:33
overlap and this is like the old radios I'm not sure whether you had that but in
11:35
I'm not sure whether you had that but in
11:35
I'm not sure whether you had that but in my childhood we used to find the FM uh
11:39
my childhood we used to find the FM uh
11:39
my childhood we used to find the FM uh channel on Sunday afternoon by tuning
11:42
channel on Sunday afternoon by tuning
11:42
channel on Sunday afternoon by tuning left and right and moving the antenna a
11:45
left and right and moving the antenna a
11:45
left and right and moving the antenna a little bit here and there so this is
11:47
little bit here and there so this is
11:47
little bit here and there so this is like that so number the chunk size and
11:49
like that so number the chunk size and
11:49
like that so number the chunk size and the chunk overlap there is no there are
11:51
the chunk overlap there is no there are
11:51
the chunk overlap there is no there are no silver bullets for this so you need
11:53
no silver bullets for this so you need
11:53
no silver bullets for this so you need to figure out depending on your document
11:56
to figure out depending on your document
11:56
to figure out depending on your document depending on your size of the document
11:58
depending on your size of the document
11:58
depending on your size of the document it can be on the information that you
12:00
it can be on the information that you
12:00
it can be on the information that you have on the document what should be the
12:01
have on the document what should be the
12:01
have on the document what should be the proper chunk size what should be the
12:03
proper chunk size what should be the
12:03
proper chunk size what should be the proper chunk overlap and you need to
12:05
proper chunk overlap and you need to
12:05
proper chunk overlap and you need to tune this properly to have your
12:08
tune this properly to have your
12:08
tune this properly to have your retrieval process also better to be
12:13
honest um there are lot of chunking um
12:17
honest um there are lot of chunking um
12:17
honest um there are lot of chunking um libraries available to be honest so if
12:19
libraries available to be honest so if
12:19
libraries available to be honest so if you are um if you are using llama index
12:21
you are um if you are using llama index
12:21
you are um if you are using llama index or L chain they have a lot of um uh
12:25
or L chain they have a lot of um uh
12:25
or L chain they have a lot of um uh libraries to help you do the chunking so
12:28
libraries to help you do the chunking so
12:28
libraries to help you do the chunking so you don't have to worry about it you
12:29
you don't have to worry about it you
12:30
you don't have to worry about it you just call a code and it chunks it um for
12:32
just call a code and it chunks it um for
12:32
just call a code and it chunks it um for you the the the most um interesting one
12:37
you the the the most um interesting one
12:37
you the the the most um interesting one that I find right now is the semantic
12:39
that I find right now is the semantic
12:39
that I find right now is the semantic splitter so you try to split and chunk
12:42
splitter so you try to split and chunk
12:42
splitter so you try to split and chunk your data depending on the semantic
12:44
your data depending on the semantic
12:44
your data depending on the semantic meaning of the data so it is using an
12:47
meaning of the data so it is using an
12:47
meaning of the data so it is using an llm to figure out what should be the
12:50
llm to figure out what should be the
12:50
llm to figure out what should be the chunk size I use and when you hear this
12:53
chunk size I use and when you hear this
12:53
chunk size I use and when you hear this you would think okay then that's the
12:55
you would think okay then that's the
12:55
you would think okay then that's the best chunker in the um in the list we
12:58
best chunker in the um in the list we
12:58
best chunker in the um in the list we have why do we need any others as I said
13:01
have why do we need any others as I said
13:01
have why do we need any others as I said before it depends on what kind of data
13:03
before it depends on what kind of data
13:03
before it depends on what kind of data you're chunking maybe you're chunking
13:04
you're chunking maybe you're chunking
13:04
you're chunking maybe you're chunking some kind of cod maybe you're chunking a
13:06
some kind of cod maybe you're chunking a
13:06
some kind of cod maybe you're chunking a structured HTML you don't need semantic
13:09
structured HTML you don't need semantic
13:09
structured HTML you don't need semantic splitter for that it only makes sense
13:12
splitter for that it only makes sense
13:12
splitter for that it only makes sense when you have same sort of information
13:15
when you have same sort of information
13:15
when you have same sort of information scattered through a lot of pages and
13:19
all I see already a question about rack
13:22
all I see already a question about rack
13:22
all I see already a question about rack fusion and all of those things I'll come
13:24
fusion and all of those things I'll come
13:24
fusion and all of those things I'll come to that maybe that's a good idea um so
13:26
to that maybe that's a good idea um so
13:26
to that maybe that's a good idea um so chunk is basically how do you do
13:28
chunk is basically how do you do
13:28
chunk is basically how do you do overlapping
13:29
overlapping
13:29
overlapping and this is a very naive uh retrieval
13:32
and this is a very naive uh retrieval
13:32
and this is a very naive uh retrieval strategy so you have documents you do a
13:34
strategy so you have documents you do a
13:34
strategy so you have documents you do a single chunk vectors then you do a
13:36
single chunk vectors then you do a
13:36
single chunk vectors then you do a search you get talk okay and you return
13:39
search you get talk okay and you return
13:39
search you get talk okay and you return now comes a different variation of um
13:44
now comes a different variation of um
13:44
now comes a different variation of um chunking so what you do is you create
13:46
chunking so what you do is you create
13:46
chunking so what you do is you create two Vector databases one has the index
13:50
two Vector databases one has the index
13:50
two Vector databases one has the index of the summary Vector so you summarize X
13:53
of the summary Vector so you summarize X
13:53
of the summary Vector so you summarize X number of pages from the documents and
13:55
number of pages from the documents and
13:55
number of pages from the documents and store it in a vector and you also have
13:58
store it in a vector and you also have
13:58
store it in a vector and you also have the normal Vector store which does the
14:00
the normal Vector store which does the
14:00
the normal Vector store which does the chunking and overlapping properly and
14:02
chunking and overlapping properly and
14:02
chunking and overlapping properly and you do a query on both of them so once
14:05
you do a query on both of them so once
14:05
you do a query on both of them so once you have the first chunk retrieve from
14:07
you have the first chunk retrieve from
14:07
you have the first chunk retrieve from the summary vectors you only retrieve
14:10
the summary vectors you only retrieve
14:10
the summary vectors you only retrieve the uh child vectors that are related to
14:13
the uh child vectors that are related to
14:14
the uh child vectors that are related to that summary so you skip the search on
14:17
that summary so you skip the search on
14:17
that summary so you skip the search on the second Vector store and that gives
14:19
the second Vector store and that gives
14:19
the second Vector store and that gives you a little bit of edge because then
14:21
you a little bit of edge because then
14:21
you a little bit of edge because then you are minimizing the search operation
14:23
you are minimizing the search operation
14:23
you are minimizing the search operation so remember the bigger your vector store
14:26
so remember the bigger your vector store
14:26
so remember the bigger your vector store is the slower your search can be
14:30
is the slower your search can be
14:30
is the slower your search can be secondly there is another thing called
14:33
secondly there is another thing called
14:33
secondly there is another thing called sentence window retrieval so for example
14:35
sentence window retrieval so for example
14:35
sentence window retrieval so for example this is a text from uh cuberes service
14:37
this is a text from uh cuberes service
14:37
this is a text from uh cuberes service automatic offer from Azure so let's say
14:40
automatic offer from Azure so let's say
14:40
automatic offer from Azure so let's say you want to find out what are the things
14:41
you want to find out what are the things
14:41
you want to find out what are the things preconfigured in Azure kubernetes and
14:45
preconfigured in Azure kubernetes and
14:45
preconfigured in Azure kubernetes and you actually should get this answer back
14:47
you actually should get this answer back
14:47
you actually should get this answer back because this is the by default um
14:50
because this is the by default um
14:50
because this is the by default um configuration you get from kubernetes
14:52
configuration you get from kubernetes
14:52
configuration you get from kubernetes but to get a better reasoning you can
14:56
but to get a better reasoning you can
14:56
but to get a better reasoning you can extend the context so you can do a
14:57
extend the context so you can do a
14:57
extend the context so you can do a sentence window retrieval so this is
14:59
sentence window retrieval so this is
14:59
sentence window retrieval so this is another way of sending the information
15:01
another way of sending the information
15:01
another way of sending the information to llm so not only the part that you
15:03
to llm so not only the part that you
15:03
to llm so not only the part that you retrieve but also the part that is
15:05
retrieve but also the part that is
15:06
retrieve but also the part that is basically around it to make a better
15:08
basically around it to make a better
15:08
basically around it to make a better judgment called better better uh
15:12
judgment called better better uh
15:12
judgment called better better uh response there's another strategy called
15:14
response there's another strategy called
15:14
response there's another strategy called parent child so you will have two
15:16
parent child so you will have two
15:16
parent child so you will have two different Vector store one vector store
15:19
different Vector store one vector store
15:19
different Vector store one vector store will have all the child leaves vectors
15:21
will have all the child leaves vectors
15:22
will have all the child leaves vectors and another will have a parent so think
15:24
and another will have a parent so think
15:24
and another will have a parent so think from this perspective if you are if your
15:26
from this perspective if you are if your
15:26
from this perspective if you are if your documents have different chapters and
15:28
documents have different chapters and
15:28
documents have different chapters and every chapter different pages so every
15:30
every chapter different pages so every
15:30
every chapter different pages so every page will have a will be considered as
15:33
page will have a will be considered as
15:33
page will have a will be considered as child and leaf and your chapter will be
15:35
child and leaf and your chapter will be
15:36
child and leaf and your chapter will be chunked as a parent and once you do a
15:38
chunked as a parent and once you do a
15:38
chunked as a parent and once you do a search on your leaf and let's say you
15:41
search on your leaf and let's say you
15:41
search on your leaf and let's say you get uh top X results back and then you
15:45
get uh top X results back and then you
15:46
get uh top X results back and then you need to judge which out of this x
15:50
need to judge which out of this x
15:51
need to judge which out of this x resolves which have the most parent
15:54
resolves which have the most parent
15:54
resolves which have the most parent relevance so you pick up that parent
15:56
relevance so you pick up that parent
15:56
relevance so you pick up that parent chapter and send it to llm and ignore
15:59
chapter and send it to llm and ignore
15:59
chapter and send it to llm and ignore the rest so that makes your answer or
16:02
the rest so that makes your answer or
16:02
the rest so that makes your answer or reasoning very much concentrated on the
16:05
reasoning very much concentrated on the
16:05
reasoning very much concentrated on the chapter where the child chunks are
16:07
chapter where the child chunks are
16:07
chapter where the child chunks are retrieve
16:10
from um there's some other thing as our
16:13
from um there's some other thing as our
16:13
from um there's some other thing as our Fusion retrieval so you will have two
16:15
Fusion retrieval so you will have two
16:15
Fusion retrieval so you will have two different Vector index one is a normal
16:17
different Vector index one is a normal
16:17
different Vector index one is a normal Vector index another is a sparse andr so
16:20
Vector index another is a sparse andr so
16:20
Vector index another is a sparse andr so this is basically using a different sort
16:22
this is basically using a different sort
16:22
this is basically using a different sort of uh vectorizing algorithm to making
16:25
of uh vectorizing algorithm to making
16:25
of uh vectorizing algorithm to making your documents vectorized and what you
16:27
your documents vectorized and what you
16:27
your documents vectorized and what you do is after retri receiving both from
16:30
do is after retri receiving both from
16:30
do is after retri receiving both from both the vector databases you do a
16:31
both the vector databases you do a
16:32
both the vector databases you do a reciprocal Fusion so you basically try
16:34
reciprocal Fusion so you basically try
16:34
reciprocal Fusion so you basically try to figure out um which all documents
16:37
to figure out um which all documents
16:37
to figure out um which all documents should be on the top of the search and
16:40
should be on the top of the search and
16:40
should be on the top of the search and which all documents should be on the
16:41
which all documents should be on the
16:41
which all documents should be on the bottom of my search results and then you
16:44
bottom of my search results and then you
16:44
bottom of my search results and then you get top X from the last results to uh
16:49
get top X from the last results to uh
16:49
get top X from the last results to uh respond to
16:50
respond to
16:51
respond to them there are a lot more to be honest I
16:54
them there are a lot more to be honest I
16:54
them there are a lot more to be honest I just picked up a few just to give you
16:56
just picked up a few just to give you
16:56
just picked up a few just to give you the feel of how uh complex it can get to
17:00
the feel of how uh complex it can get to
17:00
the feel of how uh complex it can get to be honest and there could be a lots of
17:02
be honest and there could be a lots of
17:02
be honest and there could be a lots of practical challenges doing this first of
17:04
practical challenges doing this first of
17:04
practical challenges doing this first of all um the number of um uh documents
17:08
all um the number of um uh documents
17:08
all um the number of um uh documents that you have to uh index and since you
17:11
that you have to uh index and since you
17:11
that you have to uh index and since you are doing it on a float um data type you
17:16
are doing it on a float um data type you
17:16
are doing it on a float um data type you need a lot of space so space is a major
17:19
need a lot of space so space is a major
17:19
need a lot of space so space is a major concern right now with all sorts of
17:21
concern right now with all sorts of
17:21
concern right now with all sorts of vector databases right now so they have
17:23
vector databases right now so they have
17:23
vector databases right now so they have if you have a uh 10 GB of data normally
17:27
if you have a uh 10 GB of data normally
17:27
if you have a uh 10 GB of data normally it will come around a terab when you do
17:29
it will come around a terab when you do
17:29
it will come around a terab when you do a vectorization of those data um and
17:33
a vectorization of those data um and
17:33
a vectorization of those data um and Depends also on which kind of vector
17:34
Depends also on which kind of vector
17:34
Depends also on which kind of vector mearing model you are using right now
17:36
mearing model you are using right now
17:36
mearing model you are using right now open AI gives you 1536 dimensions of
17:39
open AI gives you 1536 dimensions of
17:39
open AI gives you 1536 dimensions of vector do you need them that amount of
17:43
vector do you need them that amount of
17:43
vector do you need them that amount of dimensions for your documentation maybe
17:45
dimensions for your documentation maybe
17:45
dimensions for your documentation maybe a smaller um Dimension would also help
17:49
a smaller um Dimension would also help
17:49
a smaller um Dimension would also help so a lot of nuances are there also in
17:52
so a lot of nuances are there also in
17:52
so a lot of nuances are there also in that uh Factor when you are vectorizing
17:54
that uh Factor when you are vectorizing
17:54
that uh Factor when you are vectorizing your
17:56
your
17:56
your data um when we create a a a software
18:02
data um when we create a a a software
18:02
data um when we create a a a software architecture I would always like to
18:04
architecture I would always like to
18:04
architecture I would always like to follow a recipe like if I have a recipe
18:07
follow a recipe like if I have a recipe
18:07
follow a recipe like if I have a recipe then I know these are my steps of course
18:10
then I know these are my steps of course
18:10
then I know these are my steps of course there will be things that I can add and
18:12
there will be things that I can add and
18:12
there will be things that I can add and change and update a bit but following a
18:15
change and update a bit but following a
18:15
change and update a bit but following a recipe is always help and in rag also
18:19
recipe is always help and in rag also
18:19
recipe is always help and in rag also there is always a workflow which could
18:22
there is always a workflow which could
18:22
there is always a workflow which could be always the same depending on the
18:24
be always the same depending on the
18:24
be always the same depending on the implementation of the different boxes it
18:26
implementation of the different boxes it
18:26
implementation of the different boxes it depends how you end up with your rag
18:29
depends how you end up with your rag
18:29
depends how you end up with your rag application so for example you always
18:31
application so for example you always
18:31
application so for example you always reading from a document storage files
18:33
reading from a document storage files
18:33
reading from a document storage files and blob and you will publish some sort
18:37
and blob and you will publish some sort
18:37
and blob and you will publish some sort of document update events so for example
18:41
of document update events so for example
18:41
of document update events so for example um for example um uh um these update
18:46
um for example um uh um these update
18:46
um for example um uh um these update events will go to um somewhere a piece
18:49
events will go to um somewhere a piece
18:49
events will go to um somewhere a piece of code which will load these updates
18:52
of code which will load these updates
18:52
of code which will load these updates create embeddings or vectorization of
18:54
create embeddings or vectorization of
18:54
create embeddings or vectorization of these documents or not and then it'll
18:57
these documents or not and then it'll
18:57
these documents or not and then it'll add it to the vector
18:59
add it to the vector
18:59
add it to the vector along with some metadata um I'll part
19:01
along with some metadata um I'll part
19:01
along with some metadata um I'll part the I'll come to the metadata part in a
19:04
the I'll come to the metadata part in a
19:04
the I'll come to the metadata part in a bit and then you upsert all those
19:07
bit and then you upsert all those
19:07
bit and then you upsert all those embedding and the metadata to
19:14
somewhere
19:15
somewhere
19:15
somewhere um then comes a very interesting part
19:18
um then comes a very interesting part
19:18
um then comes a very interesting part which I think every rag application
19:20
which I think every rag application
19:21
which I think every rag application should think about from the very
19:22
should think about from the very
19:22
should think about from the very beginning is uh validation uh or testing
19:27
beginning is uh validation uh or testing
19:27
beginning is uh validation uh or testing what you call it in normal software
19:29
what you call it in normal software
19:29
what you call it in normal software development and I call it as Anon
19:31
development and I call it as Anon
19:31
development and I call it as Anon approval so from this phase you want to
19:32
approval so from this phase you want to
19:32
approval so from this phase you want to have a this phas uh expression back and
19:36
have a this phas uh expression back and
19:36
have a this phas uh expression back and how do you do it U there are a lot of
19:38
how do you do it U there are a lot of
19:38
how do you do it U there are a lot of Frameworks available which lets you
19:40
Frameworks available which lets you
19:41
Frameworks available which lets you validate your rag uh workflow the one
19:44
validate your rag uh workflow the one
19:45
validate your rag uh workflow the one that I uh want to recommend right now is
19:47
that I uh want to recommend right now is
19:47
that I uh want to recommend right now is the ragas framework uh it's from um it's
19:52
the ragas framework uh it's from um it's
19:52
the ragas framework uh it's from um it's quite quite U used framework in doing
19:55
quite quite U used framework in doing
19:55
quite quite U used framework in doing validation within um rag application and
19:59
validation within um rag application and
19:59
validation within um rag application and what does it do it try to score your rag
20:02
what does it do it try to score your rag
20:02
what does it do it try to score your rag application in four different um
20:05
application in four different um
20:05
application in four different um variables one is a context relevancy so
20:08
variables one is a context relevancy so
20:08
variables one is a context relevancy so when you are retrieving data how much
20:10
when you are retrieving data how much
20:10
when you are retrieving data how much extra unwanted information you also
20:13
extra unwanted information you also
20:13
extra unwanted information you also retrieving from the database so how much
20:16
retrieving from the database so how much
20:16
retrieving from the database so how much basically how much noise is there
20:18
basically how much noise is there
20:18
basically how much noise is there context recall um do we have all the
20:21
context recall um do we have all the
20:21
context recall um do we have all the info needed that uh was required to
20:24
info needed that uh was required to
20:25
info needed that uh was required to answer this question retrieve so your
20:27
answer this question retrieve so your
20:27
answer this question retrieve so your search result contains everything or not
20:31
search result contains everything or not
20:31
search result contains everything or not faithfulness what is the accuracy of
20:33
faithfulness what is the accuracy of
20:33
faithfulness what is the accuracy of your generated answer and answer
20:35
your generated answer and answer
20:35
your generated answer and answer relevance is how relevant is the answer
20:37
relevance is how relevant is the answer
20:38
relevance is how relevant is the answer to the query so basically you're trying
20:39
to the query so basically you're trying
20:39
to the query so basically you're trying to minimize your hallucination and
20:43
to minimize your hallucination and
20:43
to minimize your hallucination and tuning your search results with four
20:45
tuning your search results with four
20:45
tuning your search results with four different uh parameters basically and it
20:48
different uh parameters basically and it
20:48
different uh parameters basically and it it scores between 0 to one so you know
20:52
it scores between 0 to one so you know
20:52
it scores between 0 to one so you know where do you lie in all these four
20:55
where do you lie in all these four
20:55
where do you lie in all these four aspects and how does ragas do this or or
20:58
aspects and how does ragas do this or or
20:58
aspects and how does ragas do this or or any kind of validation framework does it
21:00
any kind of validation framework does it
21:00
any kind of validation framework does it is basically it asks an llm on how the
21:04
is basically it asks an llm on how the
21:04
is basically it asks an llm on how the llm is doing sounds like pretty funny
21:07
llm is doing sounds like pretty funny
21:07
llm is doing sounds like pretty funny stuff but it's not it creates a query
21:13
stuff but it's not it creates a query
21:13
stuff but it's not it creates a query and you provide a ground TR truth to
21:16
and you provide a ground TR truth to
21:16
and you provide a ground TR truth to this query and then it uses your rag
21:19
this query and then it uses your rag
21:19
this query and then it uses your rag workflow to create the answer and the
21:21
workflow to create the answer and the
21:21
workflow to create the answer and the context and then it is sent to a a large
21:25
context and then it is sent to a a large
21:25
context and then it is sent to a a large language model it can be the same large
21:27
language model it can be the same large
21:27
language model it can be the same large language model which you're using to um
21:30
language model which you're using to um
21:30
language model which you're using to um do Implement your rag application but it
21:32
do Implement your rag application but it
21:32
do Implement your rag application but it will do a scoring on all those aspects
21:34
will do a scoring on all those aspects
21:34
will do a scoring on all those aspects that I showed you last page all those
21:37
that I showed you last page all those
21:37
that I showed you last page all those four aspects and tells you what's the
21:39
four aspects and tells you what's the
21:39
four aspects and tells you what's the score between zero and
21:42
score between zero and
21:42
score between zero and one so running adding that evaluation or
21:46
one so running adding that evaluation or
21:46
one so running adding that evaluation or validation Step at the end of the
21:47
validation Step at the end of the
21:47
validation Step at the end of the workflow is also what you need to have
21:50
workflow is also what you need to have
21:50
workflow is also what you need to have when you are implementing a rag
21:54
when you are implementing a rag
21:54
when you are implementing a rag application so yeah at the end you
21:56
application so yeah at the end you
21:56
application so yeah at the end you really want a face like this when
21:58
really want a face like this when
21:58
really want a face like this when anybody is using your rag application
22:00
anybody is using your rag application
22:00
anybody is using your rag application you cannot get something which is 100%
22:03
you cannot get something which is 100%
22:03
you cannot get something which is 100% um
22:04
um
22:04
um um perfect but around 70 to 90 is quite
22:10
um perfect but around 70 to 90 is quite
22:10
um perfect but around 70 to 90 is quite good enough to achieve on a rag
22:13
good enough to achieve on a rag
22:13
good enough to achieve on a rag application provided you also have a
22:15
application provided you also have a
22:15
application provided you also have a feedback loop somehow implemented to
22:20
it um how do you master your Masterpiece
22:23
it um how do you master your Masterpiece
22:23
it um how do you master your Masterpiece is something which I call it adding
22:26
is something which I call it adding
22:26
is something which I call it adding metadata so just like any other database
22:29
metadata so just like any other database
22:29
metadata so just like any other database when you do a query the bigger your wear
22:31
when you do a query the bigger your wear
22:31
when you do a query the bigger your wear Clause is the faster your query would be
22:34
Clause is the faster your query would be
22:34
Clause is the faster your query would be so for example um right now there is a
22:37
so for example um right now there is a
22:37
so for example um right now there is a very um sparse problem of
22:41
very um sparse problem of
22:41
very um sparse problem of adding yeah roll back access to the uh
22:45
adding yeah roll back access to the uh
22:45
adding yeah roll back access to the uh Vector databases so what we can do is
22:47
Vector databases so what we can do is
22:47
Vector databases so what we can do is adding a roll column along with other
22:50
adding a roll column along with other
22:50
adding a roll column along with other columns within the vector database so
22:52
columns within the vector database so
22:52
columns within the vector database so when you're doing a query not only doing
22:54
when you're doing a query not only doing
22:54
when you're doing a query not only doing a query on that but also add the role
22:57
a query on that but also add the role
22:57
a query on that but also add the role which US is actually invoking the query
23:00
which US is actually invoking the query
23:00
which US is actually invoking the query so in that case you concentrate the
23:02
so in that case you concentrate the
23:02
so in that case you concentrate the query and your wear Clause gets tied and
23:05
query and your wear Clause gets tied and
23:05
query and your wear Clause gets tied and you have a better performing query on a
23:07
you have a better performing query on a
23:07
you have a better performing query on a database rather than anything
23:11
else
23:13
else
23:13
else um I will talk about
23:16
um I will talk about
23:16
um I will talk about two complex um Advanced rack pattern
23:20
two complex um Advanced rack pattern
23:21
two complex um Advanced rack pattern there are much many more um recently
23:24
there are much many more um recently
23:24
there are much many more um recently Microsoft also released a graph um uh
23:27
Microsoft also released a graph um uh
23:27
Microsoft also released a graph um uh rag pattern
23:29
rag pattern
23:29
rag pattern which which is basically uses a a graph
23:33
which which is basically uses a a graph
23:33
which which is basically uses a a graph SDK to create the emings on what it
23:36
SDK to create the emings on what it
23:36
SDK to create the emings on what it should uh query on so um check that out
23:41
should uh query on so um check that out
23:41
should uh query on so um check that out I'll have um Links at the end of the
23:43
I'll have um Links at the end of the
23:43
I'll have um Links at the end of the thing but the first one I want to talk
23:45
thing but the first one I want to talk
23:45
thing but the first one I want to talk about is a query transformation um what
23:47
about is a query transformation um what
23:47
about is a query transformation um what does it do it's basically if you provide
23:49
does it do it's basically if you provide
23:50
does it do it's basically if you provide a question to the llm or to the rag
23:52
a question to the llm or to the rag
23:52
a question to the llm or to the rag application it will create some
23:55
application it will create some
23:55
application it will create some subqueries from those uh from that me
23:58
subqueries from those uh from that me
23:58
subqueries from those uh from that me questions so for example if I'm trying
23:59
questions so for example if I'm trying
23:59
questions so for example if I'm trying to find out how many schools are there
24:02
to find out how many schools are there
24:02
to find out how many schools are there in New York and
24:03
in New York and
24:03
in New York and Philadelphia so it'll create two
24:05
Philadelphia so it'll create two
24:05
Philadelphia so it'll create two different queries for example how many
24:07
different queries for example how many
24:07
different queries for example how many elementary schools are there in New York
24:09
elementary schools are there in New York
24:09
elementary schools are there in New York and how many elementary schools in
24:10
and how many elementary schools in
24:10
and how many elementary schools in Philadelphia and it will do a search on
24:13
Philadelphia and it will do a search on
24:13
Philadelphia and it will do a search on both of the things and get results back
24:15
both of the things and get results back
24:15
both of the things and get results back and send all both the results back to
24:18
and send all both the results back to
24:18
and send all both the results back to llm to generate an answer so the query
24:20
llm to generate an answer so the query
24:21
llm to generate an answer so the query transformation basically gives llm a way
24:23
transformation basically gives llm a way
24:23
transformation basically gives llm a way to minimize the hallucination or make a
24:27
to minimize the hallucination or make a
24:27
to minimize the hallucination or make a retrieval much better so you are doing
24:29
retrieval much better so you are doing
24:29
retrieval much better so you are doing an um efficient query here and also an
24:34
an um efficient query here and also an
24:34
an um efficient query here and also an efficient reasoning here in LM
24:37
efficient reasoning here in LM
24:37
efficient reasoning here in LM part now I saw a question somewhere
24:39
part now I saw a question somewhere
24:39
part now I saw a question somewhere about uh mult agent multi-agent uh
24:43
about uh mult agent multi-agent uh
24:43
about uh mult agent multi-agent uh application the nuances there are a lot
24:44
application the nuances there are a lot
24:45
application the nuances there are a lot of nuances there but what is basically a
24:47
of nuances there but what is basically a
24:47
of nuances there but what is basically a multi-agent
24:48
multi-agent
24:49
multi-agent application u a multi-agent application
24:51
application u a multi-agent application
24:51
application u a multi-agent application is basically you divide your um um rag
24:56
is basically you divide your um um rag
24:56
is basically you divide your um um rag application in smaller parts so there be
24:58
application in smaller parts so there be
24:58
application in smaller parts so there be a agent orchestrator which is which has
25:01
a agent orchestrator which is which has
25:01
a agent orchestrator which is which has to be connected to an llm somehow and
25:04
to be connected to an llm somehow and
25:04
to be connected to an llm somehow and depending on the query coming in it
25:06
depending on the query coming in it
25:06
depending on the query coming in it decides which all underlying rag
25:09
decides which all underlying rag
25:09
decides which all underlying rag applications I should invol and those
25:12
applications I should invol and those
25:12
applications I should invol and those are called as an agent and it can be
25:14
are called as an agent and it can be
25:14
are called as an agent and it can be multiple agents that can be invoked or
25:16
multiple agents that can be invoked or
25:16
multiple agents that can be invoked or there can be a single agent invoked and
25:19
there can be a single agent invoked and
25:19
there can be a single agent invoked and it is also quite optional that a agent
25:21
it is also quite optional that a agent
25:21
it is also quite optional that a agent also has an llm associated with it and
25:24
also has an llm associated with it and
25:24
also has an llm associated with it and you can think on this agent you will
25:27
you can think on this agent you will
25:27
you can think on this agent you will also have a query transformation running
25:29
also have a query transformation running
25:29
also have a query transformation running there or you were doing an hybrid search
25:31
there or you were doing an hybrid search
25:31
there or you were doing an hybrid search there or you doing something more uh
25:34
there or you doing something more uh
25:34
there or you doing something more uh dedicated and depending on which what
25:37
dedicated and depending on which what
25:37
dedicated and depending on which what type of vector index you're quering and
25:39
type of vector index you're quering and
25:39
type of vector index you're quering and your answer would be different than
25:48
that last but not least there are many
25:51
that last but not least there are many
25:51
that last but not least there are many more examples but I put four major
25:56
more examples but I put four major
25:56
more examples but I put four major industries to be honest so supply chain
25:59
industries to be honest so supply chain
25:59
industries to be honest so supply chain for example if you're doing compliance
26:00
for example if you're doing compliance
26:00
for example if you're doing compliance checks or for example validation uh for
26:03
checks or for example validation uh for
26:03
checks or for example validation uh for all incoming um orders and purchase
26:07
all incoming um orders and purchase
26:07
all incoming um orders and purchase things that are going out you can
26:09
things that are going out you can
26:09
things that are going out you can automate that using a rag application um
26:12
automate that using a rag application um
26:12
automate that using a rag application um if you're doing reporting validation
26:14
if you're doing reporting validation
26:14
if you're doing reporting validation procurement B2B sales marketing supply
26:17
procurement B2B sales marketing supply
26:17
procurement B2B sales marketing supply chain has a lot of use cases on that
26:19
chain has a lot of use cases on that
26:19
chain has a lot of use cases on that part as well customer support on the
26:21
part as well customer support on the
26:21
part as well customer support on the retail is a very um well-known uh um use
26:26
retail is a very um well-known uh um use
26:26
retail is a very um well-known uh um use case to be honest so you have a lot of
26:28
case to be honest so you have a lot of
26:28
case to be honest so you have a lot of documents a new employee comes in he or
26:31
documents a new employee comes in he or
26:31
documents a new employee comes in he or she doesn't need to go through
26:32
she doesn't need to go through
26:32
she doesn't need to go through everything to understand and provide a
26:34
everything to understand and provide a
26:34
everything to understand and provide a customer proper support this person can
26:36
customer proper support this person can
26:36
customer proper support this person can query on all the datas that are there
26:38
query on all the datas that are there
26:38
query on all the datas that are there you can use product recommender for a
26:40
you can use product recommender for a
26:40
you can use product recommender for a rag you can do feedback analysis with
26:42
rag you can do feedback analysis with
26:42
rag you can do feedback analysis with rag you can do marketing uh content with
26:45
rag you can do marketing uh content with
26:45
rag you can do marketing uh content with ranks in a finance and banking case you
26:48
ranks in a finance and banking case you
26:48
ranks in a finance and banking case you can do claim processing you can do
26:50
can do claim processing you can do
26:50
can do claim processing you can do portfolio management transaction
26:52
portfolio management transaction
26:52
portfolio management transaction analytics Trend forecasting document
26:54
analytics Trend forecasting document
26:54
analytics Trend forecasting document processing lots of applications of rag
26:57
processing lots of applications of rag
26:57
processing lots of applications of rag if you are trying to achieve or
26:58
if you are trying to achieve or
26:58
if you are trying to achieve or implement it properly but it's not that
27:01
implement it properly but it's not that
27:01
implement it properly but it's not that easy there are nuances I saw some
27:03
easy there are nuances I saw some
27:03
easy there are nuances I saw some questions in there I didn't talk about
27:05
questions in there I didn't talk about
27:05
questions in there I didn't talk about all the nuances but provided you
27:08
all the nuances but provided you
27:08
all the nuances but provided you basically the the start architecture on
27:11
basically the the start architecture on
27:11
basically the the start architecture on how you can use it within your
27:13
how you can use it within your
27:13
how you can use it within your application and it's an evolving um um
27:17
application and it's an evolving um um
27:17
application and it's an evolving um um topic to be honest in last six seven
27:20
topic to be honest in last six seven
27:20
topic to be honest in last six seven months only the amount of progress um we
27:25
months only the amount of progress um we
27:25
months only the amount of progress um we had in rag um is enormous so keep an eye
27:28
had in rag um is enormous so keep an eye
27:28
had in rag um is enormous so keep an eye on this everything every day there's new
27:31
on this everything every day there's new
27:31
on this everything every day there's new thing coming up every day the um Vector
27:34
thing coming up every day the um Vector
27:34
thing coming up every day the um Vector databases are Reinventing themselves as
27:37
databases are Reinventing themselves as
27:37
databases are Reinventing themselves as I was saying a vectorization of data
27:39
I was saying a vectorization of data
27:39
I was saying a vectorization of data takes a lot of space um so all major
27:43
takes a lot of space um so all major
27:43
takes a lot of space um so all major Vector databases of quadrant um vv8 U
27:47
Vector databases of quadrant um vv8 U
27:47
Vector databases of quadrant um vv8 U Azure AI search we are also trying to uh
27:50
Azure AI search we are also trying to uh
27:50
Azure AI search we are also trying to uh quantize this data so binary
27:52
quantize this data so binary
27:52
quantize this data so binary quantization will make the floating
27:54
quantization will make the floating
27:54
quantization will make the floating Point data to an integer data and it
27:57
Point data to an integer data and it
27:57
Point data to an integer data and it will make the storage space uh smaller
28:00
will make the storage space uh smaller
28:00
will make the storage space uh smaller than previously so it will also save
28:03
than previously so it will also save
28:03
than previously so it will also save your money when you're operating on a
28:05
your money when you're operating on a
28:05
your money when you're operating on a large scale rag there are a lot of
28:07
large scale rag there are a lot of
28:07
large scale rag there are a lot of things happening within rag um area so
28:09
things happening within rag um area so
28:09
things happening within rag um area so keep an eye on it and um last but not
28:13
keep an eye on it and um last but not
28:13
keep an eye on it and um last but not least since this is a kind of a do
28:16
least since this is a kind of a do
28:16
least since this is a kind of a do community I wanted to talk about a
28:18
community I wanted to talk about a
28:18
community I wanted to talk about a little bit when you use semantic kernel
28:22
little bit when you use semantic kernel
28:22
little bit when you use semantic kernel to implement a r
28:24
to implement a r
28:24
to implement a r application and what would the flow look
28:26
application and what would the flow look
28:27
application and what would the flow look like so when a question comes in
28:28
like so when a question comes in
28:28
like so when a question comes in semantic will um vectorize the question
28:32
semantic will um vectorize the question
28:32
semantic will um vectorize the question do a query on the vector
28:34
do a query on the vector
28:34
do a query on the vector store and then possibly can also uh
28:38
store and then possibly can also uh
28:39
store and then possibly can also uh invoke external apis or um user database
28:43
invoke external apis or um user database
28:43
invoke external apis or um user database or some sort of other functions to
28:45
or some sort of other functions to
28:45
or some sort of other functions to gather more information and then using
28:48
gather more information and then using
28:48
gather more information and then using all those context from Vector databases
28:51
all those context from Vector databases
28:51
all those context from Vector databases and from the external API it'll create
28:54
and from the external API it'll create
28:54
and from the external API it'll create an answer and send it back and this is
28:57
an answer and send it back and this is
28:57
an answer and send it back and this is all possible with this is the semantic
29:00
all possible with this is the semantic
29:00
all possible with this is the semantic kernel but you can also do it with llama
29:01
kernel but you can also do it with llama
29:01
kernel but you can also do it with llama index you can also do with llama index
29:03
index you can also do with llama index
29:03
index you can also do with llama index and also with a land chain to be honest
29:06
and also with a land chain to be honest
29:06
and also with a land chain to be honest and semantic kernel has a net SDK which
29:09
and semantic kernel has a net SDK which
29:09
and semantic kernel has a net SDK which basically quite nice nobody else in
29:13
basically quite nice nobody else in
29:13
basically quite nice nobody else in the um AI orchestration area has a net
29:18
the um AI orchestration area has a net
29:18
the um AI orchestration area has a net um
29:20
um
29:20
um SDK with that I always prefer that uh
29:25
SDK with that I always prefer that uh
29:25
SDK with that I always prefer that uh not only um what I have a knowledge to
29:28
not only um what I have a knowledge to
29:28
not only um what I have a knowledge to share but also sharing how you learn
29:31
share but also sharing how you learn
29:31
share but also sharing how you learn more so there are a few articles which I
29:34
more so there are a few articles which I
29:34
more so there are a few articles which I quite liked when I talked about complex
29:37
quite liked when I talked about complex
29:37
quite liked when I talked about complex rack and there are a couple of examples
29:39
rack and there are a couple of examples
29:39
rack and there are a couple of examples from Azure side which talks about how
29:41
from Azure side which talks about how
29:41
from Azure side which talks about how you implement graph Rack or how you
29:43
you implement graph Rack or how you
29:43
you implement graph Rack or how you implement a normal open AI chat GPT
29:46
implement a normal open AI chat GPT
29:46
implement a normal open AI chat GPT style um rag application using uh Azure
29:53
style um rag application using uh Azure
29:53
style um rag application using uh Azure resources if you have question
29:59
[Music]


