Practical FinOps: Optimizing Generative AI Costs in GCP Vertex AI

May 02, 20266 min read
"FinOps""GCP""Vertex AI""Cloud"

With the massive adoption of Language Models (LLMs) in business processes, public cloud bills (AWS, GCP, Azure) began to inflate rapidly due to the cost of token-based API calls (input/output). The role of the infrastructure professional is to enable technological innovation while ensuring strict financial efficiency (FinOps).

In this article, I show how I structured financial boundaries and governance for the AI team to test prompts and develop agents in GCP Vertex AI without exceeding the budget.

1. Token Cost Monitoring

We implemented consolidated billing filters in GCP Billing, divided by key-value labeling (labels) linked to each project. In this way, we were able to measure exactly which AI model (such as Gemini 1.5 Pro or Flash) was consuming the most financial resources in the RAG pipeline.

2. Quotas and Request Limits

Instead of leaving chat APIs open without restrictions in development testing environments, we applied strict quotas for requests per minute (RPM) and requests per day (RPD) within the GCP Quota Manager.

3. Gemini Flash vs Pro

One of the most efficient FinOps actions was to direct data pre-processing and simple classification workflows to the Gemini Flash model, leaving the more expensive Gemini Pro focused exclusively on the final synthesis and reasoning phase of complex RAG.

Quota and label commands

Day-to-day control is not a console click. Quota and billing labels show which model blew the Opex.

gcloud services quota list \
  --service=aiplatform.googleapis.com \
  --consumer=projects/PROJECT_ID

gcloud billing budgets list --billing-account=BILLING_ACCOUNT_ID

The label rides on the request so billing can split Flash from Pro:

curl -s -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  "https://us-central1-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/us-central1/publishers/google/models/gemini-2.0-flash:generateContent" \
  -d '{"contents":[{"role":"user","parts":[{"text":"Classify this ticket: disk at 90%."}]}],"labels":{"workload":"triage"}}'

Classification and preprocessing stay on Flash. RAG synthesis stays on Pro.


Conclusion

This division of model responsibility and quota control reduced monthly AI development cost projections by over 60%, demonstrating that technological agility can and should go hand-in-hand with financial health (Opex).

Related Articles