Building a Local RAG Agent with Python, Ollama (Llama 3), and CPU
Introduction
Nowadays, information security is one of the biggest concerns for companies, especially when dealing with generative artificial intelligence. Sending confidential corporate data (such as audit reports, ISO compliance policies, or financial statements) to external proprietary APIs can violate strict data governance policies.
The solution? 100% Local RAG (Retrieval-Augmented Generation).
In this article, we will build a local RAG agent step-by-step that runs entirely on your machine (even using only CPU) using Python, Ollama (with the Llama 3 model), and LangChain/ChromaDB.
What is RAG (Retrieval-Augmented Generation)?
RAG is an architecture that extends the capabilities of a Large Language Model (LLM) without the need to retrain it (fine-tuning). It works in three main steps:
- Retrieval: The system searches for relevant documents from a local knowledge base based on the user's question.
- Augmentation: The user's question and the found document fragments are combined into a "rich prompt."
- Generation: The LLM reads the provided context and generates an accurate response, based solely on the content of the provided documents.
Requirements and Environment Setup
1. Install Ollama
Access ollama.com and install the software on your system (Linux, macOS, or Windows). After installation, open the terminal and download the Llama 3.1 model (or Llama 3 8B):ollama pull llama3
ollama pull nomic-embed-text
nomic-embed-text is a lightweight, high-performance embedding model to convert text into mathematical vectors.
2. Python Dependencies
Create a virtual environment and install the dependencies:python -m venv venv
source venv/bin/activate
pip install langchain langchain-community chromadb sentence-transformers langchain-text-splitters
The Local RAG Agent Code
Below is the complete script that reads local documents (such as PDFs or text files), creates the local vector database (ChromaDB), and executes local inference using Ollama.
import os
from langchain_community.document_loaders import TextLoader
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_community.embeddings import OllamaEmbeddings
from langchain_community.vectorstores import Chroma
from langchain_community.llms import Ollama
from langchain.chains import RetrievalQA
from langchain.prompts import PromptTemplate
1. Load the local knowledge document
Suppose we have a 'manual_sgi.txt' file with the company rules
with open("manual_sgi.txt", "w", encoding="utf-8") as f:
f.write("""
MANUAL DO SISTEMA DE GESTÃO INTEGRADO (SGI) - ALEX FERRAZ TECNOLOGIA
Seção 4.2 - Política de Segurança de Redes corporativas:
Todos os servidores na AWS devem usar autenticação via chave SSH pública/privada (RSA 4096).
É proibido o acesso root direto via SSH. Apenas o usuário 'deploy' tem acesso ao ambiente de produção.
As senhas de banco de dados devem ser rotacionadas a cada 30 dias usando o AWS Secrets Manager.
""")
loader = TextLoader("manual_sgi.txt", encoding="utf-8")
documents = loader.load()
2. Partition the text into smaller blocks (Chunking)
text_splitter = RecursiveCharacterTextSplitter(chunk_size=300, chunk_overlap=50)
texts = text_splitter.split_documents(documents)
3. Generate Embeddings and create the Vector Database (Chroma DB)
Chroma will persist the data in the ./chroma_db folder locally and securely
embeddings = OllamaEmbeddings(model="nomic-embed-text")
vector_store = Chroma.from_documents(
documents=texts,
embedding=embeddings,
persist_directory="./chroma_db"
)
4. Create the Custom Prompt to Guarantee Scope Limitation
template = """Use as seguintes partes do contexto fornecido para responder à pergunta no final.
Se você não souber a resposta, diga apenas que não sabe, não tente inventar uma resposta.
Mantenha a resposta curta, direta e estritamente baseada no contexto.
Contexto:
{context}
Pergunta: {question}
Resposta útil em português:"""
custom_prompt = PromptTemplate(
template=template,
input_variables=["context", "question"]
)
5. Configure the Llama3 Model in Ollama and the Question & Answer Chain
llm = Ollama(model="llama3", temperature=0.1)
qa_chain = RetrievalQA.from_chain_type(
llm=llm,
chain_type="stuff",
retriever=vector_store.as_retriever(search_kwargs={"k": 2}),
chain_type_kwargs={"prompt": custom_prompt}
)
6. Execute the Question
pergunta = "Qual usuário tem acesso ao ambiente de produção e qual a política de chaves SSH?"
resposta = qa_chain.run(pergunta)
print("--- PERGUNTA ---")
print(pergunta)
print("\n--- RESPOSTA DO AGENTE ---")
print(resposta)
CPU Performance Analysis
Embeddings: The nomic-embed-text model generates embeddings almost instantaneously on the CPU, as it has few parameters.
Semantic Search: ChromaDB is extremely lightweight and retrieves snippets in milliseconds.
* Generation (Llama 3 8B): On a modern CPU (e.g., AMD Ryzen 7 or Intel Core i7), generation takes about 10 to 20 seconds per response. Using 4-bit quantization (Ollama default), RAM consumption is around 4.8 GB.
Commands before the script
The Python below assumes Ollama is already answering. These three commands are the check I run before starting the agent, including on a CPU-only machine.
ollama pull llama3
ollama pull nomic-embed-text
ollama list
curl -s http://127.0.0.1:11434/api/tags
api/tags must list llama3 on 127.0.0.1. If the port shows up on 0.0.0.0, the model is exposed on the network and the ISO 27001 argument (and the SOC 2 access-control argument) falls apart.
Conclusion and Governance
This architecture demonstrates that it is feasible to build robust corporate AI assistants without relying on the internet or exposing industrial secrets. For systems requiring ISO 27001 (Information Security) compliance, local RAG definitively eliminates the risk of data leakage.
Did you like the tutorial? If you are planning to implement local Generative AI in your infrastructure, share your questions on my LinkedIn!
Related Articles
How to Implement Generative AI and Local RAG in Compliance with ISO 27001 and ISO 9001
Local RAG with Ollama on CPU, aligned with ISO 27001, ISO 9001, and the access control a SOC 2 audit also asks for.
Multi-cloud for Banks: Separating the Core, Data, and the Audit Trail
Multi-cloud architecture for banks: a locked region, a core isolated from the lab, and an audit trail nobody can delete. SCP, CloudTrail, and Object Lock commands.