Building a Local RAG Agent with Python, Ollama (Llama 3), and CPU

August 01, 20268 min read
PythonAIRAGOllamaLangChain

Introduction

Nowadays, information security is one of the biggest concerns for companies, especially when dealing with generative artificial intelligence. Sending confidential corporate data (such as audit reports, ISO compliance policies, or financial statements) to external proprietary APIs can violate strict data governance policies.

The solution? 100% Local RAG (Retrieval-Augmented Generation).

In this article, we will build a local RAG agent step-by-step that runs entirely on your machine (even using only CPU) using Python, Ollama (with the Llama 3 model), and LangChain/ChromaDB.


What is RAG (Retrieval-Augmented Generation)?

RAG is an architecture that extends the capabilities of a Large Language Model (LLM) without the need to retrain it (fine-tuning). It works in three main steps:

  1. Retrieval: The system searches for relevant documents from a local knowledge base based on the user's question.
  2. Augmentation: The user's question and the found document fragments are combined into a "rich prompt."
  3. Generation: The LLM reads the provided context and generates an accurate response, based solely on the content of the provided documents.

Requirements and Environment Setup

1. Install Ollama

Access ollama.com and install the software on your system (Linux, macOS, or Windows). After installation, open the terminal and download the Llama 3.1 model (or Llama 3 8B):
ollama pull llama3
ollama pull nomic-embed-text
nomic-embed-text is a lightweight, high-performance embedding model to convert text into mathematical vectors.

2. Python Dependencies

Create a virtual environment and install the dependencies:
python -m venv venv
source venv/bin/activate
pip install langchain langchain-community chromadb sentence-transformers langchain-text-splitters

The Local RAG Agent Code

Below is the complete script that reads local documents (such as PDFs or text files), creates the local vector database (ChromaDB), and executes local inference using Ollama.

import os
from langchain_community.document_loaders import TextLoader
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_community.embeddings import OllamaEmbeddings
from langchain_community.vectorstores import Chroma
from langchain_community.llms import Ollama
from langchain.chains import RetrievalQA
from langchain.prompts import PromptTemplate

1. Load the local knowledge document

Suppose we have a 'manual_sgi.txt' file with the company rules

with open("manual_sgi.txt", "w", encoding="utf-8") as f: f.write(""" MANUAL DO SISTEMA DE GESTÃO INTEGRADO (SGI) - ALEX FERRAZ TECNOLOGIA Seção 4.2 - Política de Segurança de Redes corporativas: Todos os servidores na AWS devem usar autenticação via chave SSH pública/privada (RSA 4096). É proibido o acesso root direto via SSH. Apenas o usuário 'deploy' tem acesso ao ambiente de produção. As senhas de banco de dados devem ser rotacionadas a cada 30 dias usando o AWS Secrets Manager. """)

loader = TextLoader("manual_sgi.txt", encoding="utf-8") documents = loader.load()

2. Partition the text into smaller blocks (Chunking)

text_splitter = RecursiveCharacterTextSplitter(chunk_size=300, chunk_overlap=50) texts = text_splitter.split_documents(documents)

3. Generate Embeddings and create the Vector Database (Chroma DB)

Chroma will persist the data in the ./chroma_db folder locally and securely

embeddings = OllamaEmbeddings(model="nomic-embed-text") vector_store = Chroma.from_documents( documents=texts, embedding=embeddings, persist_directory="./chroma_db" )

4. Create the Custom Prompt to Guarantee Scope Limitation

template = """Use as seguintes partes do contexto fornecido para responder à pergunta no final. Se você não souber a resposta, diga apenas que não sabe, não tente inventar uma resposta. Mantenha a resposta curta, direta e estritamente baseada no contexto.

Contexto: {context}

Pergunta: {question} Resposta útil em português:"""

custom_prompt = PromptTemplate( template=template, input_variables=["context", "question"] )

5. Configure the Llama3 Model in Ollama and the Question & Answer Chain

llm = Ollama(model="llama3", temperature=0.1)

qa_chain = RetrievalQA.from_chain_type( llm=llm, chain_type="stuff", retriever=vector_store.as_retriever(search_kwargs={"k": 2}), chain_type_kwargs={"prompt": custom_prompt} )

6. Execute the Question

pergunta = "Qual usuário tem acesso ao ambiente de produção e qual a política de chaves SSH?" resposta = qa_chain.run(pergunta)

print("--- PERGUNTA ---") print(pergunta) print("\n--- RESPOSTA DO AGENTE ---") print(resposta)


CPU Performance Analysis

Embeddings: The nomic-embed-text model generates embeddings almost instantaneously on the CPU, as it has few parameters. Semantic Search: ChromaDB is extremely lightweight and retrieves snippets in milliseconds. * Generation (Llama 3 8B): On a modern CPU (e.g., AMD Ryzen 7 or Intel Core i7), generation takes about 10 to 20 seconds per response. Using 4-bit quantization (Ollama default), RAM consumption is around 4.8 GB.


Commands before the script

The Python below assumes Ollama is already answering. These three commands are the check I run before starting the agent, including on a CPU-only machine.

ollama pull llama3
ollama pull nomic-embed-text
ollama list
curl -s http://127.0.0.1:11434/api/tags

api/tags must list llama3 on 127.0.0.1. If the port shows up on 0.0.0.0, the model is exposed on the network and the ISO 27001 argument (and the SOC 2 access-control argument) falls apart.


Conclusion and Governance

This architecture demonstrates that it is feasible to build robust corporate AI assistants without relying on the internet or exposing industrial secrets. For systems requiring ISO 27001 (Information Security) compliance, local RAG definitively eliminates the risk of data leakage.

Did you like the tutorial? If you are planning to implement local Generative AI in your infrastructure, share your questions on my LinkedIn!

Related Articles