Skip to main content

Command Palette

Search for a command to run...

보이는 챗봇 구현을 위한 멀티모달 Rag 시스템 구축하기

멀티-모달 RAG 챗봇을 LangChain과 GPT-4o를 사용하여 구축하고, 질답 기능을 구현해 보세요.

Published
•View as Markdown
보이는 챗봇 구현을 위한 멀티모달 Rag 시스템 구축하기

원문: Bhargob Deka, "Building a Multi-Modal RAG System for Visual Question Answering"

멀티모달 RAG 챗봇용 Streamlit 앱

개요

이 글에서는 OpenAI의 GPT-4o 모델을 사용하여 멀티모달 RAG 채팅 애플리케이션을 구축하는 과정을 안내합니다. 다음과 같은 내용을 배우게 될 것입니다.

  • 멀티모달 RAG 채팅 애플리케이션: PDF 문서에서 정보를 검색하여 시각적 질문 응답을 가능하게 하는 애플리케이션을 만듭니다.

  • 매끄러운 파싱: Unstructured 라이브러리를 사용하여 텍스트, 표, 이미지 등을 매끄럽게 파싱하는 방법을 배웁니다.

  • 성능 평가: DeepEval 라이브러리를 통해 다양한 지표를 사용하여 챗봇의 성능을 평가합니다.

  • Streamlit UI: Streamlit 앱을 통해 애플리케이션을 시연합니다.

이 글을 읽어야 하는 이유

고급 기반 모델인 GPT-4o의 멀티모달 기능을 활용하여 AI 애플리케이션을 구축하고 싶으신가요? 그렇다면 이 글이 딱 맞는 자료입니다!

마케팅 전문가로서 시장 조사 보고서에서 인사이트를 찾거나, 의료 전문가로서 멀티모달 의료 문서를 분석하거나, 법률 전문가로서 복잡한 법률 파일을 처리하고자 하는 분들께 이 글은 유용한 통찰이 될 것입니다.

각 개념을 자세히 설명하고 코드에 대한 상세한 설명을 제공하겠습니다. 그럼 시작해 볼까요! 🎬

멀티모달 RAG의 부상

텍스트 기반 RAG 모델에서 멀티모달 RAG 시스템으로의 전환은 AI 능력에 큰 발전이 있었다는 것을 의미합니다. 간단한 개요를 살펴보겠습니다.

  • 유래: RAG라는 용어는 2021년 4월에 처음 등장했으며, 원래는 텍스트 기반 지식을 통해 언어 출력을 향상하는 방식이었습니다.

  • 발전: 2024년 5월에 출시된 GPT-4o와 같은 모델을 통해 이제는 시각적 정보를 통합하여 이미지, 표, 텍스트를 동시에 처리할 수 있습니다.

  • 새로운 가능성: 이 진화는 더 포괄적이고 맥락적으로 풍부한 AI 애플리케이션을 가능하게 합니다.

이 글에서는 제 연구 논문 중 하나인 Neurocomputing에서 멀티모달 RAG 프레임워크를 사용한 사례 연구를 알려 드립니다. 논문에는 텍스트, 표, 그래프가 포함되어 있으며, GPT-4o의 시각적 기능을 활용하여 복잡한 질문에 어떻게 답할 수 있는지 탐구해 보겠습니다.

그럼 코딩을 시작해 볼까요! 🎬

목차

  1. 가상 환경 설정 및 파이썬 라이브러리 설치

  2. 비정형 데이터 전처리

  3. 텍스트, 표, 이미지 요약

  4. 멀티모달 검색기

  5. 멀티모달 RAG 체인

  6. LLM 평가

  7. Streamlit을 이용한 사용자 인터페이스

가상 환경 설정 및 파이썬 라이브러리 설치

먼저 가상 환경을 설정해 봅시다. 다음 명령어를 사용하세요.

python3.10 -m venv venv

이제 필요한 패키지들을 설치해 봅시다. 해당 패키지들은 GitHub 저장소의 메인 디렉터리에 있는 'requirements.txt' 파일에서 찾을 수 있습니다.

pip install -r requirements.txt

이제 "your-project.ipynb"라는 Jupyter 노트북을 열어 코드를 작성해 봅시다. 이제 준비는 끝났습니다! 이제 본격적으로 주요 내용을 다뤄보겠습니다.

아래는 멀티모달 RAG 프로젝트를 위한 GitHub 저장소 링크입니다.

비정형 데이터 전처리

RAG 애플리케이션을 구축하기 위한 첫 번째 단계는, PDF 문서를 사용하여 데이터를 데이터베이스에 로드하는 것입니다. 대형 언어 모델(LLM)의 콘텍스트 창 제한 때문에, 문서 전체를 직접 프롬프트에 넣고 저장할 수는 없습니다. 그렇게 하면 최대 토큰 수를 초과하여 오류가 발생할 가능성이 높기 때문입니다.

이를 해결하기 위해, 문서에서 이미지, 텍스트, 표와 같은 다양한 요소를 추출하는 작업을 시작할 것입니다. 이 작업에는 Unstructured 라이브러리를 사용할 것입니다.

설치

아직 pip로 설치하지 않았다면 먼저 패키지를 설치해 보겠습니다.

#%brew install tesseract poppler
%pip install -q "unstructured[all-docs]"

참고로, Unstructured 라이브러리가 텍스트 추출 및 이미지에서 텍스트를 추출할 수 있도록 시스템에 'tesseract'와 'poppler' 라이브러리도 필요합니다. 두 패키지는 Homebrew를 사용하여 설치할 수 있습니다 (주석 처리된 명령어를 참고하세요).

파티셔닝 및 청킹

from unstructured.partition.pdf import partition_pdf

elements = partition_pdf(
    filename="TAGIV.pdf", # mandatory
    strategy="hi_res",                                     # mandatory to use ``hi_res`` strategy
    extract_images_in_pdf=True,                            # mandatory to set as ``True``
    extract_image_block_types=["Image", "Table"],          # optional
    extract_image_block_to_payload=False,                  # optional
    extract_image_block_output_dir="saved_images",  # optional - only works when ``extract_image_block_to_payload=False``
    )

우리는 partition_pdf 모듈을 사용하여 문서를 분할하고, 파일 'TAGIV.pdf'에서 다양한 요소를 추출할 것입니다. 고해상도 이미지와 표를 추출하기 위해 hi_res 전략을 설정할 것입니다. 선택적인 매개변수 extract_image_block_types와 extract_image_block_output_dir를 사용하여 이미지와 표만 추출하고, 이를 "saved_images"라는 디렉터리에 저장하도록 지정할 수 있습니다.

이제 추출된 요소들을 chunk_by_title 방법을 사용하여 청킹할 것입니다. 보통 서론, 방법론, 결과 등 별도의 섹션과 하위 섹션으로 구성되는 연구 논문에 적합합니다.

from unstructured.chunking.title import chunk_by_title # might be better for an article
from typing import Any

chunks = chunk_by_title(elements)

# different category in the document
category_counts = {}

for element in chunks:
   category = str(type(element))
   if category in category_counts:
       category_counts[category] += 1
   else:
       category_counts[category] = 1

# Unique_categories will have unique elements
unique_categories = set(category_counts.keys())
category_counts
{"<class 'unstructured.documents.elements.CompositeElement'>": 200,
"<class 'unstructured.documents.elements.Table'>": 3,
"<class 'unstructured.documents.elements.TableChunk'>": 2}

청킹 결과 다음의 세 가지 고유한 카테고리가 나타났습니다.

  • CompositeElements

  • Table

  • TableChunk

'CompositeElements'는 단락, 섹션, 푸터, 수식 등 다양한 텍스트 요소들의 모음입니다. 또한 세 개의 'Table' 구조와 두 개의 'TableChunk'가 있으며, 'TableChunk'는 보통 표의 일부 또는 세그먼트를 나타냅니다. 즉, 표가 여러 페이지에 걸쳐 나뉘어 있고 그중 일부만 청킹된 경우일 수 있습니다.

문서에는 네 개의 표가 있는데, 그중 세 개만 완전히 파싱되었습니다. 🤔

필터링

다음 단계로, 문서 요소들을 간소화하여 텍스트와 표 데이터를 각각 별도로 처리할 수 있도록 할 것입니다. 이를 위해 Pydantic 모델을 정의하여 문서 요소를 표준화하고, 해당 요소들을 유형에 따라 "텍스트" 또는 "표"로 분류할 것입니다.

from pydantic import BaseModel

class Element(BaseModel):
   type: str
   text: Any


# Categorize by type
categorized_elements = []
for element in chunks:
   if "unstructured.documents.elements.CompositeElement" in str(type(element)):
       categorized_elements.append(Element(type="text", text=str(element)))
   elif "unstructured.documents.elements.Table" in str(type(element)):
       categorized_elements.append(Element(type="table", text=str(element)))

# Text
text_elements = [e for e in categorized_elements if e.type == "text"]

# Table
table_elements = [e for e in categorized_elements if e.type == "table"]

우리는 문서 요소 청크를 순회하며 각 요소의 유형을 식별한 후, 이를 분류된 리스트에 추가할 것입니다. 마지막으로 이 리스트를 텍스트와 표 요소로 나누어 별도의 리스트로 필터링합니다. 이로써 전처리 단계가 완료됩니다.

텍스트, 표, 이미지 요약

이후 멀티벡터 검색기를 사용할 준비를 하기 위해, 텍스트, 표, 이미지 요소의 요약을 만들어야 합니다. 이러한 요약들은 벡터 저장소에 저장되어, 입력 쿼리를 프롬프트로 전달할 때 시맨틱 검색이 가능하게 해줍니다.

텍스트와 표 요약

우선 텍스트와 표 요약부터 시작해보겠습니다. 먼저 AI에게 테이블과 텍스트를 요약하는 연구 보조 역할을 하도록 지시하는 프롬프트 템플릿을 설정합니다. 그다음, 이 프롬프트와 GPT-4o 모델을 통해 각 텍스트 및 표 요소를 처리하여 간결한 요약을 생성하는 체인을 만듭니다.

효율성을 위해, 매개변수 max_concurrency를 사용하여 다섯 개의 텍스트 또는 표 요소를 동시에 배치 처리할 것입니다.

%pip install -q langchain langchain-chroma unstructured[all-docs] pydantic lxml langchainhub langchain-openai

from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_openai import ChatOpenAI

## Retriever

# Prompt
prompt_text = """You are an expert Research Assistant tasked with summarizing tables and texts from research articles. \
Give a concise summary of the text. text chunk: {element} """

prompt = ChatPromptTemplate.from_template(prompt_text)

# Summary chain
model = ChatOpenAI(temperature=0, model="gpt-4o")
summarize_chain = {"element": lambda x: x} | prompt | model | StrOutputParser()

# Apply to texts
texts = [i.text for i in text_elements]
text_summaries = summarize_chain.batch(texts, {"max_concurrency": 5})

# Apply to tables
tables = [i.text for i in table_elements]
table_summaries = summarize_chain.batch(tables, {"max_concurrency": 5})

이미지 요약

다음으로, 이미지를 요약하는 데 도움을 줄 함수들을 만들겠습니다. encode_image, image_summarize, generate_img_summaries의 세 가지 주요 함수를 정의할 것입니다.

  1. encode_image: 이 함수는 이미지를 바이너리 읽기 모드(‘rb’)로 열어 base64로 인코딩된 문자열을 반환합니다.

  2. image_summarize: 이 함수는 모델에게 이미지를 요약하는 방법을 지시하는 프롬프트를 포함하는 HumanMessage 객체를 사용합니다. 또한, base64로 인코딩된 이미지 데이터를 데이터 URL 형식으로 포함하여 콘텐츠에 이미지를 직접 삽입할 수 있도록 합니다.

  3. generate_img_summaries: 이 함수는 지정된 디렉터리 내의 모든 JPG 이미지를 처리하여 각 이미지에 대한 요약을 생성하고, base64로 인코딩된 이미지들을 반환합니다.

이 함수들을 통해 이미지를 효율적으로 요약하고 처리할 수 있으며, 이를 멀티모달 RAG 애플리케이션에 원활하게 통합할 수 있습니다.

다음은 전체 코드입니다.

## getting image summaries
import base64
import os

from langchain_core.messages import HumanMessage


def encode_image(image_path):
   """Getting the base64 string"""
   with open(image_path, "rb") as image_file:
       return base64.b64encode(image_file.read()).decode("utf-8")

def image_summarize(img_base64, prompt):
   """Make image summary"""
   chat = ChatOpenAI(model="gpt-4o", max_tokens=1024)

   msg = chat.invoke(
       [
           HumanMessage(
               content=[
                   {"type": "text", "text": prompt},
                   {
                       "type": "image_url",
                       "image_url": {"url": f"data:image/jpg;base64,{img_base64}"},
                   },
               ]
           )
       ]
   )
   return msg.content


def generate_img_summaries(path):
   """
   Generate summaries and base64 encoded strings for images
   path: Path to list of .jpg files extracted by Unstructured
   """

   # Store base64 encoded images
   img_base64_list = []

   # Store image summaries
   image_summaries = []

   # Prompt
   prompt = """You are an assistant tasked with summarizing images for retrieval. \
   These summaries will be embedded and used to retrieve the raw image. \
   Give a concise summary of the image that is well optimized for retrieval."""

   # Apply to images
   for img_file in sorted(os.listdir(path)):
       if img_file.endswith(".jpg"):
           img_path = os.path.join(path, img_file)
           base64_image = encode_image(img_path)
           img_base64_list.append(base64_image)
           image_summaries.append(image_summarize(base64_image, prompt))

   return img_base64_list, image_summaries


fpath = "saved_images"

# Image summaries
img_base64_list, image_summaries = generate_img_summaries(fpath)

멀티모달 검색기

요약이 준비되었으니, 이제 멀티모달 검색기를 만들어봅시다.

멀티벡터 검색기

우리는 vectorstore, docstore, id_key, search_kwargs를 입력으로 받는 멀티벡터 검색기를 설정할 것입니다. 이 방법은 콘텐츠 요약을 별도로 인덱싱하면서 원본 콘텐츠를 저장할 수 있어 효율적인 검색이 가능하게 합니다. 이것은 멀티모달 RAG를 수행하는 하나의 방법일 뿐이며, 또 다른 방법으로는 CLIP을 사용하여 텍스트와 이미지를 멀티모달 임베딩으로 변환한 다음, 원본 이미지와 텍스트 청크를 멀티모달 LLM에 전달하는 방식이 있을 수 있습니다. 이 방식은 나중에 블로그에서 다룰 예정입니다. 🙂

우리의 검색기는 Chroma 벡터스토어를 활용해 콘텐츠 요약의 임베딩을 저장하고, InMemoryStore를 사용해 전체 콘텐츠를 저장합니다. 이 설정은 요약을 통해 시맨틱 검색을 수행하면서, 필요한 경우 해당 요약에 맞는 원본 콘텐츠를 검색할 수 있게 해 줍니다. 각 문서에는 고유 식별자가 UUID를 통해 할당되며, 이는 검색기에서 필수적으로 사용됩니다.

또한 요약을 벡터스토어에 추가하고 원본 콘텐츠를 docstore에 추가하는 과정을 간소화하기 위해 add_documents라는 헬퍼 함수를 만들 것입니다. 이 함수는 사용할 수 있는 요약만 추가되도록 보장합니다.

import uuid

from langchain.retrievers.multi_vector import MultiVectorRetriever
from langchain.storage import InMemoryStore
from langchain_chroma import Chroma
from langchain_core.documents import Document
from langchain_openai import OpenAIEmbeddings

def create_multi_vector_retriever(
   vectorstore, text_summaries, texts, table_summaries, tables, image_summaries, images):
   """
   Create retriever that indexes summaries, but returns raw images, table, or texts
   """

   # Initialize the storage layer
   store = InMemoryStore()
   id_key = "doc_id"

   # Create the multi-vector retriever
   retriever = MultiVectorRetriever(
       vectorstore=vectorstore,
       docstore=store,
       id_key=id_key,
       search_kwargs={"k": 2}  # Limit to top 2 results
   )

   # Helper function to add documents to the vectorstore and docstore
   def add_documents(retriever, doc_summaries, doc_contents):
       doc_ids = [str(uuid.uuid4()) for  in doccontents]
       summary_docs = [
           Document(page_content=s, metadata={id_key: doc_ids[i]})
           for i, s in enumerate(doc_summaries)
       ]
       retriever.vectorstore.add_documents(summary_docs)
       retriever.docstore.mset(list(zip(doc_ids, doc_contents)))

   # Add texts, tables, and images
   # Check that text_summaries is not empty before adding
   if text_summaries:
       add_documents(retriever, text_summaries, texts)
   # # Check that table_summaries is not empty before adding
   if table_summaries:
       add_documents(retriever, table_summaries, tables)
   # Check that image_summaries is not empty before adding
   if image_summaries:
       add_documents(retriever, image_summaries, images)

   return retriever

검색기 생성

이제 OpenAI 임베딩 모델을 사용하는 Chroma 벡터스토어를 할당하고, 검색기를 생성해 봅시다.

# The vectorstore to use to index the summaries
vectorstore = Chroma(
   collection_name="mm_tagiv_paper", embedding_function=OpenAIEmbeddings()
)

# Create retriever
retriever_multi_vector_img = create_multi_vector_retriever(
   vectorstore,
   text_summaries,
   texts,
   table_summaries,
   tables,
   image_summaries,
   img_base64_list,
)

테스트

이제 이 쿼리를 사용해 검색기를 테스트하여 어떤 문서들이 검색되는지 확인해 봅시다.

retriever_multi_vector_img.invoke("How is the performance of TAGI-V for the Boston dataset compared to the other methods?")
['TAGI-V are averaged over 3 random seeds. The test log-likelihood values show that TAGI-V performs better than all other methods in 4 out of the 5 datasets. The TAGI-V method is also competitive for RMSE values where it provides the best results in 2 out of the 5 datasets, i.e., Elevators and KeggD, while it is second best for KeggU and Pol. Both PCA+ VI and NL outperform the others in two datasets.']

응답을 확인해 보니 문서에서 특정 정보를 성공적으로 찾아냈습니다. 완벽하군요! 🚀

멀티모달 RAG 체인

이제 검색기가 준비되었으니, 멀티모달 체인을 만들어보겠습니다. base64로 인코딩된 이미지와 텍스트 데이터를 처리하기 위해 몇 가지 헬퍼 함수가 필요합니다.

헬퍼 함수

  • plt_image_base64(img_base64): base64로 인코딩된 이미지를 HTML을 사용해 표시하는 함수입니다.
def plt_img_base64(img_base64):
   image_html = f'<img src="data:image/jpg;base64,{img_base64}" />'
   display(HTML(image_html))
  • looks_like_base64(sb): 문자열이 base64로 인코딩된 것인지 확인하는 함수입니다.
def looks_like_base64(sb):
   return re.match("^[A-Za-z0-9+/]+[=]{0,2}$", sb) is not None
  • is_image_data(b64data): base64 데이터의 헤더를 확인하여 해당 데이터가 이미지를 나타내는지 검증하는 함수입니다.
def is_image_data(b64data):
   image_signatures = {
       b"\xff\xd8\xff": "jpg",
       b"\x89\x50\x4e\x47\x0d\x0a\x1a\x0a": "png",
       b"\x47\x49\x46\x38": "gif",
       b"\x52\x49\x46\x46": "webp",
   }
   try:
       header = base64.b64decode(b64data)[:8]
       for sig, format in image_signatures.items():
           if header.startswith(sig):
               return True
       return False
   except Exception:
       return False
  • resize_base64_image(base64_string, size=(128, 128)): base64로 인코딩된 이미지를 지정된 크기로 조정하는 함수입니다.

      def resize_base64_image(base64_string, size=(128, 128)):
         img_data = base64.b64decode(base64_string)
         img = Image.open(io.BytesIO(img_data))
         resized_img = img.resize(size, Image.LANCZOS)
         buffered = io.BytesIO()
         resized_img.save(buffered, format=img.format)
         return base64.b64encode(buffered.getvalue()).decode("utf-8")
    
  • split_image_text_types(docs): 문서 리스트를 base64로 인코딩된 이미지와 텍스트로 분리하는 함수입니다.

def split_image_text_types(docs):
   b64_images = []
   texts = []
   for doc in docs:
       if isinstance(doc, Document):
           doc = doc.page_content
       if looks_like_base64(doc) and is_image_data(doc):
           doc = resize_base64_image(doc, size=(1300, 600))
           b64_images.append(doc)
       else:
           texts.append(doc)
   return {"images": b64_images, "texts": texts}

프롬프트 함수

함수 img_prompt_func(data_dict)는 AI 모델을 위한 입력 데이터를 포맷팅하는 함수입니다. 이 함수는 텍스트와 이미지 데이터를 하나의 프롬프트로 결합합니다. 여기에는 '사용자 질문'과 '채팅 기록'이 포함됩니다.

def img_prompt_func(data_dict):
   formatted_texts = "\n".join(data_dict["context"]["texts"])
   messages = []

   if data_dict["context"]["images"]:
       for image in data_dict["context"]["images"]:
           image_message = {
               "type": "image_url",
               "image_url": {"url": f"data:image/jpg;base64,{image}"},
           }
           messages.append(image_message)

   chat_history = data_dict.get("chat_history", [])
   formatted_chat_history = "\n".join([f"{m.type}: {m.content}" for m in chat_history])

   text_message = {
       "type": "text",
       "text": (
           "You are a Research Assistant tasked with answering questions on research articles.\n"
           "You will be given a mixed of text, tables, and image(s) usually of tables, charts or graphs.\n"
           "Use this information to provide accurate information related to the user question. \n"
           f"User-provided question: {data_dict['question']}\n\n"
           "Text and / or tables:\n"
           f"{formatted_texts}"
           "Chat History:\n"
           f"{formatted_chat_history}\n\n"
       ),
   }
   messages.append(text_message)
   return [HumanMessage(content=messages)]

멀티모달 RAG 체인

마지막으로 multi_modal_rag_chain(retriever, memory=None) 함수는 RAG 체인을 설정하는 데 사용됩니다. 체인의 작동 방식은 다음과 같습니다.

  • 먼저 RunnableParallel 컴포넌트가 시작합니다. 이 컴포넌트는 관련 문서를 동시에 검색하고, split_image_text_types 함수를 사용하여 문서를 텍스트와 이미지로 분리합니다. 동시에 사용자의 질문을 그대로 전달하고, 메모리에서 대화 기록을 가져옵니다. 이러한 병렬 처리는 필요한 모든 맥락 정보를 빠르고 효과적으로 수집하는 것을 보장합니다.

  • 이 단계의 출력은 img_prompt_func에 의해 구조화된 프롬프트로 포맷팅되며, 이는 사용자 쿼리, 검색된 맥락, 그리고 채팅 기록을 AI 모델이 처리하기 적합한 형식으로 통합합니다.

  • 이 구조화된 프롬프트는 GPT-4o 모델에 전달되어 제공된 정보를 바탕으로 응답을 생성합니다.

  • 마지막으로 StrOutputParser가 모델의 출력을 문자열로 포맷팅하여, 이후 사용할 준비를 완료합니다.

이 설계는 시스템이 텍스트와 시각적 데이터를 모두 이해하고 통합해야 하는 복잡한 쿼리를 능숙하게 처리하면서, 진행 중인 대화의 맥락을 유지할 수 있도록 합니다.

def multi_modal_rag_chain(retriever, memory=None):
   if memory is None:
       memory = ConversationBufferMemory(return_messages=True, memory_key="chat_history")

   model = ChatOpenAI(temperature=0, model="gpt-4o", max_tokens=1024)

   chain = (
       RunnableParallel(
           {
           "context": retriever | RunnableLambda(split_image_text_types),
           "question": RunnablePassthrough(),
           "chat_history": lambda x: memory.load_memory_variables({})["chat_history"]
       })
       | RunnableLambda(img_prompt_func)
       | model
       | StrOutputParser()
   )

   def run_chain(query):
       result = chain.invoke(query)
       memory.save_context({"input": query}, {"output": result})
       return result

   return run_chain


# Create RAG chain
chain_mm_rag = multi_modal_rag_chain(retriever=retriever_multi_vector_img)

테스트 시간!

이제 첫 번째 질문을 던져 보고, 체인의 응답을 확인해 봅시다.

# First Question
query = "How is the performance of TAGI-V for the Boston dataset compared to the other methods?"
print(chain_mm_rag(query))
To determine the performance of TAGI-V for the Boston dataset compared to other methods, we need to look at the specific metrics provided for this dataset. The text mentions that TAGI-V performs better than all other methods in 4 out of the 5 datasets for test log-likelihood values and is competitive for RMSE values, providing the best results in 2 out of the 5 datasets.

However, the text does not explicitly state the performance of TAGI-V on the Boston dataset. To provide a precise answer, we would need the specific test log-likelihood and RMSE values for the Boston dataset for TAGI-V and the other methods.

Given the information provided:
- TAGI-V is generally strong in test log-likelihood values.
- TAGI-V is competitive in RMSE values, being the best in some datasets and second best in others.

If the Boston dataset is one of the datasets where TAGI-V is not the best, it might be outperformed by PCA+VI or NL, which are mentioned as top performers in two datasets each.

Without the specific values for the Boston dataset, we can infer that TAGI-V is likely competitive but may not be the top performer for this particular dataset. For a definitive comparison, the exact test log-likelihood and RMSE values for the Boston dataset across all methods would be necessary.

응답이 정확합니다. 제가 논문의 저자기 때문에 확신할 수 있습니다. 😀

이제 두 번째 질문을 시도해 봅시다.

# Second Question
query = "What is the performance of the same method for the Concrete dataset compared to the other methods?"
print(chain_mm_rag(query))
To assess the performance of the same method for the Concrete dataset compared to other methods, we can refer to the provided graphs and tables. Here is a detailed analysis:

### Performance Metrics:
1. RMSE (Root Mean Square Error):
  - The RMSE values for the Concrete dataset are depicted in the graphs. The methods compared include PCA+ESS, PCA+VI, SWAG, TAGI-V, TAGI-V2L, TAGI, PBP, MC-dropout, PBP-MV, VMG, Ensemble, DVI, and NN.
  - From the graph, it appears that TAGI-V and its variants (TAGI-V2L, TAGI) have competitive RMSE values compared to other methods. The exact RMSE values are not explicitly provided in the text, but the visual representation indicates that TAGI-V performs well.

2. Training Time:
  - The training time for the Concrete dataset is shown in the time (s) vs. RMSE graph.
  - TAGI-V and its variants (TAGI-V2L, TAGI) demonstrate faster training times compared to methods like PCA+ESS, PCA+VI, PBP-MV, and VMG. TAGI-V is significantly faster, approximately 100 times faster than PCA+ESS and PCA+VI, about 10 times faster than PBP, and about 3 times faster than Ensemble.

### Comparative Analysis:
- TAGI-V:
 - RMSE: TAGI-V shows competitive RMSE values, indicating good predictive performance.
 - Training Time: TAGI-V is significantly faster in training time compared to most other methods.

- Other Methods:
 - PCA+ESS and PCA+VI: These methods have higher RMSE values and longer training times compared to TAGI-V.
 - SWAG, PBP, MC-dropout, PBP-MV, VMG, Ensemble, DVI, NN: These methods also show higher RMSE values and longer training times compared to TAGI-V.

### Summary:
TAGI-V demonstrates superior performance for the Concrete dataset in terms of both RMSE and training time. It achieves competitive RMSE values and significantly faster training times compared to other methods. This makes TAGI-V a highly efficient and effective method for the Concrete dataset.

따라서 우리가 TAGI-V에 대해 질문하고 있다는 것을 기억하고 있으며, 이제 애플리케이션이 대화형이라는 것을 보여 줍니다. 😎

이제 base64로 인코딩된 이미지를 반환해야 하는 질문에 대해 검색된 문서를 확인해 봅시다.

# Check retrieval
query = "How is the performance of He compared to modified He for the various datasets such as Boston, Concrete etc.?"
docs = retriever_multi_vector_img.invoke(query, limit=6)

검색된 문서 중 하나가 실제로 이미지 파일로 등장했습니다. 🙌🏼

LLM 평가

모델을 평가하기 위해 LLM 평가를 위한 오픈소스 프레임워크인 DeepEval을 사용할 것입니다. 이 프레임워크는 입력 쿼리에 대해 검색된 문서와 최종 응답을 테스트할 수 있는 여러 지표를 제공합니다. 이번 실험에서는 다음과 같은 지표에 집중할 것입니다.

  • 신뢰성(Faithfulness) 지표: 모델의 출력이 제공된 맥락과 얼마나 일치하는지를 측정합니다.

  • 맥락적 관련성(Contextual Relevancy) 지표: 검색된 맥락이 주어진 쿼리와 얼마나 관련이 있는지를 평가합니다.

  • 응답 관련성(Answer Relevancy) 지표: 모델의 응답이 입력된 쿼리와 얼마나 관련이 있는지를 평가합니다.

  • 환각(Hallucination) 지표: 모델의 출력이 제공된 맥락에 없는 정보를 포함하고 있는지 감지합니다.

DeepEval은 이 외에도 다양한 지표를 제공합니다. 라이브러리 문서에서 더 많은 지표를 탐색해 보시길 권장합니다.

이제 지표 각각에 대한 기능을 포함한 LLM_Metric이라는 클래스를 정의할 것입니다. 각 지표의 출력은 단순한 점수뿐만 아니라 그 점수에 대한 이유도 제공하여, 모델 성능에 대한 더 깊은 통찰을 가져다줍니다.

class LLM_Metric:
   def init(self, query, retrieval_context, actual_output):
       self.query = query
       self.retrieval_context = retrieval_context
       self.actual_output = actual_output

   # Faithfulness
   def get_faithfulness_metric(self):
       metric = FaithfulnessMetric(
           threshold=0.7,
           model="gpt-4o",
           include_reason=True
       )
       test_case = LLMTestCase(
           input=self.query,
           actual_output=self.actual_output,
           retrieval_context=self.retrieval_context
       )

       metric.measure(test_case)
       return metric.score, metric.reason

   # Contextual Relevancy
   def get_contextual_relevancy_metric(self):
       metric = ContextualRelevancyMetric(
           threshold=0.7,
           model="gpt-4o",
           include_reason=True
       )
       test_case = LLMTestCase(
           input=self.query,
           actual_output=self.actual_output,
           retrieval_context=self.retrieval_context
       )

       metric.measure(test_case)
       return metric.score, metric.reason

   # Answer Relevancy
   def get_answer_relevancy_metric(self):
       metric = AnswerRelevancyMetric(
       threshold=0.7,
       model="gpt-4o",
       include_reason=True
       )
       test_case = LLMTestCase(
           input=self.query,
           actual_output=self.actual_output
       )
       metric.measure(test_case)
       return metric.score, metric.reason

   # Hallucination
   def get_hallucination_metric(self):
       metric = HallucinationMetric(threshold=0.5)
       test_case = LLMTestCase(
       input=self.query,
       actual_output=self.actual_output,
       context=self.retrieval_context 
       )
       metric.measure(test_case)
       return metric.score, metric.reason

Streamlit을 이용한 사용자 인터페이스

우선 코드를 모듈화된 방식으로 구조화합니다. 모든 함수를 utils라는 폴더에 넣을 것입니다. 디렉터리 구조는 다음과 같습니다.

advanced-RAG-app/
│
├── utils/
│ ├── init.py
│ ├── image_processing.py
│ ├── rag_chain.py
│ ├── rag_evaluation.py
│ └── retriever.py
│
└── main.py
└── requirements.txt

아래 함수들은 이미 앞서 설명했습니다. 여기서는 단순히 구조를 변경하여 메인 애플리케이션 파일에서 모든 함수를 호출할 수 있도록 정리하는 것입니다.

그럼 이제! main.py 파일을 작성해 봅시다.

import streamlit as st
from unstructured.partition.pdf import partition_pdf
from unstructured.chunking.title import chunk_by_title
from typing import Any
from pydantic import BaseModel
from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_openai import ChatOpenAI
from langchain_chroma import Chroma
from langchain_openai import OpenAIEmbeddings
from utils.image_processing import generate_img_summaries
from utils.retriever import create_multi_vector_retriever
from utils.rag_chain import multi_modal_rag_chain, plt_img_base64
from utils.rag_evaluation import LLM_Metric
from io import BytesIO
import base64
from PIL import Image
import io

# Initialize session state
if 'processed' not in st.session_state:
   st.session_state.processed = False
if 'retriever' not in st.session_state:
   st.session_state.retriever = None
if 'chain' not in st.session_state:
   st.session_state.chain = None

# Streamlit app setup
st.set_page_config(page_title='Multi-Modal RAG Application', page_icon='random', layout='wide', initial_sidebar_state='auto')

def process_document(uploaded_file):
   # Process PDF
   with st.spinner('Processing PDF...'):
       st.sidebar.info('Extracting elements from PDF...')
       pdf_bytes = uploaded_file.read()
       elements = partition_pdf(
           file=BytesIO(pdf_bytes),
           strategy="hi_res",
           extract_images_in_pdf=True,
           extract_image_block_types=["Image", "Table"],
           extract_image_block_to_payload=False,
           extract_image_block_output_dir="docs/saved_images",
       )
       st.sidebar.success('PDF elements extracted successfully!')

   # Create chunks by title
   with st.spinner('Chunking content...'):
       st.sidebar.info('Creating chunks by title...')
       chunks = chunk_by_title(elements)
       st.sidebar.success('Chunking completed successfully!')

   # Categorize Elements
   class Element(BaseModel):
       type: str
       text: Any

   categorized_elements = []
   for element in chunks:
       if "unstructured.documents.elements.CompositeElement" in str(type(element)):
           categorized_elements.append(Element(type="text", text=str(element)))
       elif "unstructured.documents.elements.Table" in str(type(element)):
           categorized_elements.append(Element(type="table", text=str(element)))

   text_elements = [e for e in categorized_elements if e.type == "text"]
   table_elements = [e for e in categorized_elements if e.type == "table"]

   # Prompt
   prompt_text = """You are an expert Research Assistant tasked with summarizing tables and texts from research articles. \
   Give a concise summary of the text. text chunk: {element} """

   prompt = ChatPromptTemplate.from_template(prompt_text)

   # Summary chain
   model = ChatOpenAI(temperature=0, model="gpt-4o", max_tokens=1024)
   summarize_chain = {"element": lambda x: x} | prompt | model | StrOutputParser()

   texts = [i.text for i in text_elements]
   text_summaries = summarize_chain.batch(texts, {"max_concurrency": 5})

   tables = [i.text for i in table_elements]
   table_summaries = summarize_chain.batch(tables, {"max_concurrency": 5})

   # Image summaries
   fpath = "docs/saved_images"
   img_base64_list, image_summaries = generate_img_summaries(fpath)

   # Vectorstore
   vectorstore = Chroma(
       collection_name="mm_tagiv_paper", embedding_function=OpenAIEmbeddings()
   )

   # Create retriever
   st.session_state.retriever = create_multi_vector_retriever(
       vectorstore,
       text_summaries,
       texts,
       table_summaries,
       tables,
       image_summaries,
       img_base64_list,
   )

   # Create RAG chain
   st.session_state.chain = multi_modal_rag_chain(retriever=st.session_state.retriever)
   st.session_state.processed = True

with st.sidebar:
   # File upload
   st.subheader('Add your PDF')
   uploaded_file = st.file_uploader("Upload a PDF file", type=["pdf"])
   if st.button('Submit'):
       if uploaded_file is not None:
           process_document(uploaded_file)
           st.success('Document processed successfully!')
       else:
           st.error('Please upload a PDF file first.')

# Main page for query response and evaluation
st.subheader("RAG Assistant")
query = st.text_input("Enter your query:")

if query and st.session_state.processed:
   # Execution
   retrieval_context = st.session_state.retriever.invoke(query, limit=1)
   actual_output = st.session_state.chain(query)

   # Evaluation
   llm_metric = LLM_Metric(query, retrieval_context, actual_output)
   faith_score, faith_reason = llm_metric.get_faithfulness_metric()
   relevancy_score, relevancy_reason = llm_metric.get_contextual_relevancy_metric()
   answer_relevancy_score, answer_relevancy_reason = llm_metric.get_answer_relevancy_metric()
   hallucination_score, hallucination_reason = llm_metric.get_hallucination_metric()

   # Display results
   st.subheader("Query Response")
   st.write(actual_output)

   st.subheader("Evaluation Metrics")
   st.write(f"Faithfulness Score: {faith_score}, Reason: {faith_reason}")
   st.write(f"Contextual Relevancy Score: {relevancy_score}, Reason: {relevancy_reason}")
   st.write(f"Answer Relevancy Score: {answer_relevancy_score}, Reason: {answer_relevancy_reason}")
   st.write(f"Hallucination Score: {hallucination_score}, Reason: {hallucination_reason}")


elif query and not st.session_state.processed:
   st.warning("Please upload and process a document first.")

너무 어려워하지 마세요! 제가 각 단계를 자세히 안내해 드리겠습니다.

먼저 애플리케이션은 세션 상태를 초기화하여 문서가 처리되었는지 여부를 추적하고, 검색기와 RAG 체인 객체를 저장합니다.

다음으로 process_document 함수를 정의하는데 이 함수는 업로드된 PDF 파일의 핵심 처리를 담당합니다. 여기에는 다음 작업이 포함됩니다.

  • PDF 추출

  • 청킹

  • 분류

  • 요약

  • 이미지 요약

  • 벡터스토어 초기화

  • 검색기 생성

  • RAG 체인 생성

사이드바에서는 사용자가 PDF 파일을 업로드할 수 있으며, 이때 문서 처리 함수가 실행됩니다. 문서가 성공적으로 처리되면, 메인 페이지에서 사용자가 쿼리를 입력하고, 응답과 평가 지표를 확인할 수 있습니다.

마지막으로 앱을 실행하는 명령은 다음과 같습니다.

streamlit run main.py

마침내!

마무리하며

이제 프로젝트의 마무리까지 도달했습니다. 우리는 PDF 문서에 대한 시각적 질문 응답을 위한 멀티모달 RAG 애플리케이션을 만드는 방법을 탐구했습니다. 다음과 같은 내용을 다뤘습니다.

  • Unstructured 라이브러리를 사용해 문서를 청크로 나눴습니다.

  • 멀티벡터 검색기를 사용해 텍스트, 표, 이미지 요약을 벡터 스토어에 저장하고, 원본 콘텐츠는 문서 스토어에 보관했습니다.

  • DeepEval 라이브러리를 활용해 LLM 응답을 평가했습니다.

  • Streamlit을 사용해 애플리케이션의 간단한 UI를 만들었습니다.

이 글에 대한 여러분의 생각을 댓글로 듣고 싶습니다. 여러분께 즐거운 주말 프로젝트가 되었길 바랍니다. 그럼 다음에 또 만나요!

À Bientôt 🙂


안녕하세요, 저는 Bhargob입니다! 👋

저는 생성형 AI(GenAI) 애플리케이션 개발에 큰 관심을 가진 머신러닝 연구원입니다.

제 작업을 지켜봐 주신 모든 분들께 진심으로 감사드립니다. 제 글이 여러분의 AI 학습 여정에 영감을 되었기를 바랍니다. LinkedIn과 GitHub에서 소통할 수 있으며, 흥미로운 사람들과 흥미로운 생성형 AI 프로젝트에 대한 협업을 언제든 환영합니다. 사업용 AI 솔루션*이 필요하시다면, LinkedIn을 통해 연락해 주세요.*

함께 이야기할 수 있길 기대합니다!


자료

생성형 AI 개념에 익숙하지 않으시다면, LangChain 프레임워크에 관한 제 초보자 시리즈를 따라 보시길 강력히 권장합니다. 기본 개념부터 고급 배포 전략까지, AI 애플리케이션을 직접 구축하는 데 필요한 모든 내용을 다룹니다.

다음은 우리가 다룬 내용을 간단히 요약한 것입니다.

1️⃣ LangChain을 사용한 챗봇*: 파이썬을 사용한 스무스한 입문*
2️⃣ 검색 체인 만들기*: LangChain으로 RAG 챗봇 구축*
3️⃣ LangChain을 사용한 RAG 에이전트 만들기
4️⃣ LangGraph로 RAG 에이전트 워크플로 디자인하기
5️⃣ LangGraph, FastAPI, Streamlit/Gradio*: AI 개발을 위한 완벽한 삼총사*
6️⃣ 로컬에서 클라우드로*: Docker와 AWS EC2를 사용한 LLM 애플리케이션 배포*
7️⃣ Git Push로 클라우드 배포*: GitHub Actions만 있으면 충분합니다*

고급 AI 주제에 대한 글도 준비했습니다. 두 가지 엔드 투 엔드 프로젝트로 멀티 에이전트 프레임워크를 다뤘습니다.

8️⃣ 호텔 추천 시스템 구축*: CrewAI, Ollama, Gradio를 사용한 멀티 에이전트 프레임워크*
9️⃣ 엔드 투 엔드 데이터 과학 프로젝트 생성*: CrewAI 에이전트를 활용한 프로젝트*

이 자료들이 여러분의 AI 여정에 도움이 되기를 바랍니다. 즐거운 학습되세요!

More from this blog

나의 오픈 소스 시작 이야기

원문: TkDoDo, “My Open Source Origin Story“ 가끔씩 제가 받는 질문이 하나 있는데, 바로 오픈 소스와 리액트 쿼리(React Query)를 어떻게 시작하게 되었는지입니다. 저의 기본 원칙은 어떤 질문을 세 번 받으면 더 이상 답변할 필요가 없도록 질문에 대해 글로 쓴다는 것입니다. 하지만 이 질문은 주로 직접 만났을 때 받는 질문이라 글로 작성할 생각을 한 적이 없었습니다. 최근에 오프라인 컨퍼런스에 더 많이 참...

Jul 30, 2025
나의 오픈 소스 시작 이야기

이더넷이란?

원문: baeldung, “What Is Ethernet?“ 1. 소개 이 튜토리얼에서는 이더넷(Ethernet)과 이를 통해 이루어지는 데이터 전송에 대해 알아보겠습니다. 2. 이더넷이란? 이더넷은 근거리 통신망(LAN) 또는 광역 네트워크(WAN) 내에서 장치들이 데이터를 주고받고 통신하기 쉽게 만들어 주는 널리 사용되는 기술입니다. 컴퓨터, 프린터, 서버는 물론 스마트 홈 기기까지도 이더넷으로 연결됩니다. 가정이나 사무실처럼 제한된 공간...

Jul 20, 2025
이더넷이란?

포스트 개발자 시대

원문: Josh W. Comeau, "The Post-Developer Era" 2년 전 2023년 3월, "프런트엔드 개발의 종말"이라는 제목의 블로그 글을 발행했습니다. 이는 OpenAI가 GPT-4 쇼케이스를 발표한 직후였고, 당시 업계 분위기는 머지않아 인간 소프트웨어 개발자는 필요 없어지고 앞으로는 소프트웨어 개발을 AI가 전담하게 될 것이라는 전망이 지배적이었습니다. 저는 이런 주장에 회의적이었고 그 블로그 글에서 소프트웨어 개발...

Jul 10, 2025
포스트 개발자 시대

널리 사용되는 네트워크 프로토콜

원문: Subham Datta, "Popular Network Protocols" 1. 개요 이 튜토리얼에서는 가장 널리 사용되고 인기 있는 네트워크 프로토콜들을 소개합니다. 2. 네트워크 프로토콜 소개 의사소통과 정보 교환은 현대 사회에서 가장 중요하고 강력한 역량입니다. 컴퓨터 네트워킹이란 여러 대의 컴퓨터와 장치를 케이블이나 위성을 통해 서로 연결하여, 거리와 상관없이 정보·자원·데이터베이스 등을 공유할 수 있게 하는 것을 말합니다. 네...

Jun 20, 2025
널리 사용되는 네트워크 프로토콜

커맨드 라인에 편해지는 법

원문: Julia Evans, "What helps people get comfortable on the command line?" 가끔 커맨드 라인을 써야 하는 친구들과 이야기하다 보면 많은 이들이 여전히 터미널을 두려워하고 있다는 걸 느낍니다. 그럴 때마다 어떤 조언을 할지 잘 모르겠더라고요. 저는 워낙 오래전부터 터미널을 써왔기 때문이죠. 그래서 Mastodon에 이렇게 물어봤습니다. 최근 1~3년 사이에 터미널 공포(?)를 극복한 분...

Jun 10, 2025
커맨드 라인에 편해지는 법
C

CodeSnap

84 posts

한국어로 전달하는 웹 개발 번역 매거진