Practical FinOps: Optimizing Generative AI Costs in GCP Vertex AI
With the massive adoption of Language Models (LLMs) in business processes, public cloud bills (AWS, GCP, Azure) began to inflate rapidly due to the cost of token-based API calls (input/output). The role of the infrastructure professional is to enable technological innovation while ensuring strict financial efficiency (FinOps).
In this article, I show how I structured financial boundaries and governance for the AI team to test prompts and develop agents in GCP Vertex AI without exceeding the budget.
1. Token Cost Monitoring
We implemented consolidated billing filters in GCP Billing, divided by key-value labeling (labels) linked to each project. In this way, we were able to measure exactly which AI model (such as Gemini 1.5 Pro or Flash) was consuming the most financial resources in the RAG pipeline.
2. Quotas and Request Limits
Instead of leaving chat APIs open without restrictions in development testing environments, we applied strict quotas for requests per minute (RPM) and requests per day (RPD) within the GCP Quota Manager.3. Gemini Flash vs Pro
One of the most efficient FinOps actions was to direct data pre-processing and simple classification workflows to the Gemini Flash model, leaving the more expensive Gemini Pro focused exclusively on the final synthesis and reasoning phase of complex RAG.Quota and label commands
Day-to-day control is not a console click. Quota and billing labels show which model blew the Opex.
gcloud services quota list \
--service=aiplatform.googleapis.com \
--consumer=projects/PROJECT_ID
gcloud billing budgets list --billing-account=BILLING_ACCOUNT_ID
The label rides on the request so billing can split Flash from Pro:
curl -s -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
"https://us-central1-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/us-central1/publishers/google/models/gemini-2.0-flash:generateContent" \
-d '{"contents":[{"role":"user","parts":[{"text":"Classify this ticket: disk at 90%."}]}],"labels":{"workload":"triage"}}'
Classification and preprocessing stay on Flash. RAG synthesis stays on Pro.
Conclusion
This division of model responsibility and quota control reduced monthly AI development cost projections by over 60%, demonstrating that technological agility can and should go hand-in-hand with financial health (Opex).Related Articles
Migration from On-Premises Infrastructure to Microsoft 365: Strategy and Compliance
A guide on how to plan and execute email and local file migrations from traditional Windows/Exchange servers to Microsoft 365, ensuring integrity and ISO compliance.
Multi-cloud for Banks: Separating the Core, Data, and the Audit Trail
Multi-cloud architecture for banks: a locked region, a core isolated from the lab, and an audit trail nobody can delete. SCP, CloudTrail, and Object Lock commands.