Per-model token attribution, not one GCP line
CostMon joins your GCP billing export with Cloud Monitoring's per-model token metrics, so Vertex AI spend breaks down by model and by day instead of disappearing into the rest of your GCP bill.
Vertex AI / Gemini cost monitoring
CostMon reads Vertex AI spend from the same GCP billing export as the rest of your Google Cloud bill, and joins it with Cloud Monitoring's per-model token counts, so a workload that should be running through Batch, or one that's crossed a long-context surcharge threshold, shows up in days instead of blended into next month's GCP invoice.
Vertex AI spend surfaces through the same GCP Billing Export as every other Google Cloud service, which means it's easy for it to disappear into an undifferentiated "GCP" total unless someone splits it out by service. Even then, the export alone doesn't show which model or which workload is driving it. Two of the biggest levers, running non-realtime workloads through the discounted Batch API instead of standard serving, and using context caching for repeated long prompts, are opt-in decisions that nothing in the billing console flags when they're being missed.
Cost anatomy
Why CostMon
CostMon joins your GCP billing export with Cloud Monitoring's per-model token metrics, so Vertex AI spend breaks down by model and by day instead of disappearing into the rest of your GCP bill.
A normalized daily view surfaces a high-volume, non-realtime workload running at standard rates, or a repeated-context pattern that isn't using context caching. Those are the two easiest discounts to leave on the table.
Vertex AI spend lands in the same normalized daily table as Bedrock, Anthropic, OpenAI, and every other connected provider, so finance sees one honest AI total instead of a line item buried in a cloud bill.
CostMon connects with a service account scoped to bigquery.dataViewer, jobUser, and monitoring.viewer only. It's read-only and never touches your Vertex AI models, prompts, or infrastructure.
Getting started
Grant a service account bigquery.dataViewer and jobUser on your billing export dataset, plus monitoring.viewer for token metrics. CostMon never requests access to your Vertex AI models or infrastructure.
Vertex AI rows are filtered out of your GCP billing export and joined to Cloud Monitoring's per-model input and output token counts on the same day, reconciled into the same normalized model as every other provider you connect.
Open a daily, per-model view of Vertex AI spend and catch a missed Batch discount or an uncached repeated-context workload weeks before it lands on a bill.
FAQ
Through the same read-only service account as the GCP connector (bigquery.dataViewer and jobUser on your billing export dataset), plus monitoring.viewer scoped to Cloud Monitoring's token-usage metrics. CostMon never requests access to your Vertex AI models or prompts.
No. It's the same GCP Billing Export everything else reads from. CostMon filters it to Vertex AI / generative AI services and joins it with per-model token metrics, so it's broken out without needing a separate connector or invoice.
CostMon surfaces per-model daily spend and volume, which is the signal you need to spot a high-volume, non-realtime workload running at standard online rates. It doesn't switch the workload to Batch for you, but it makes the opportunity visible.
Cached and standard input tokens both flow through Cloud Monitoring's token metrics, so a workload that isn't benefiting from context caching shows up as a higher effective input-token cost per call relative to similar workloads that are.
A flat monthly rate per plan, not a percentage of your Vertex AI spend or the savings CostMon helps you find. Free includes 2 connectors; Professional and Enterprise include unlimited connectors.
Connectors
AWS, GCP, and Azure for cloud; Anthropic, OpenAI, Amazon Bedrock, Vertex AI / Gemini, and Helicone for AI and LLM spend; Snowflake and Databricks for the data cloud; Datadog, GitHub, and Vercel for the rest of the stack. Add CSV import for any tool without a native connector. That's 14 connectors, all normalized into the same unified view.
Connect your first provider and CostMon normalizes it alongside everything else you run. Your team gets one number everyone can check.