CostMonStart free

Vertex AI / Gemini cost monitoring

See what your Gemini usage costs, per model

CostMon reads Vertex AI spend from the same GCP billing export as the rest of your Google Cloud bill, and joins it with Cloud Monitoring's per-model token counts, so a workload that should be running through Batch, or one that's crossed a long-context surcharge threshold, shows up in days instead of blended into next month's GCP invoice.

GCP billing export + Cloud Monitoring token usage — live today

Vertex AI spend surfaces through the same GCP Billing Export as every other Google Cloud service, which means it's easy for it to disappear into an undifferentiated "GCP" total unless someone splits it out by service. Even then, the export alone doesn't show which model or which workload is driving it. Two of the biggest levers, running non-realtime workloads through the discounted Batch API instead of standard serving, and using context caching for repeated long prompts, are opt-in decisions that nothing in the billing console flags when they're being missed.

Cost anatomy

What actually drives your Vertex AI bill

Per-token, priced by model tier
Vertex AI / Gemini bills per input and output token, separately, and the rate varies widely by model tier. The cheapest and most capable Gemini models can differ several-fold on the same million tokens.
Long-context surcharge
Requests above certain context-length thresholds (e.g. 200K input tokens on some models) bill at a surcharge over the standard token rate. A workload that grows its prompts over time can cross that line without anyone noticing.
Batch vs. online serving
Batch/Flex prediction is priced at roughly half the cost of standard online serving for the same model. A non-realtime workload, like bulk classification or nightly summarization, run online instead of batch leaves that discount on the table.
Context caching, its own line item
Reusing a long system prompt or knowledge-base context via context caching bills cached input tokens far below the standard input rate, plus a separate cache-storage charge. Skip it, and a repeated-context workload pays full input price on every call.
Same billing export as the rest of GCP
Vertex spend surfaces through the same GCP Billing Export in BigQuery as every other Google Cloud service, so it's easy for it to blend into an undifferentiated "GCP" total unless it's split out by service.

Why CostMon

Built for Vertex AI spend, from day one

Per-model token attribution, not one GCP line

CostMon joins your GCP billing export with Cloud Monitoring's per-model token metrics, so Vertex AI spend breaks down by model and by day instead of disappearing into the rest of your GCP bill.

Catch a missed Batch discount or caching gap

A normalized daily view surfaces a high-volume, non-realtime workload running at standard rates, or a repeated-context pattern that isn't using context caching. Those are the two easiest discounts to leave on the table.

One view, next to the rest of your AI stack

Vertex AI spend lands in the same normalized daily table as Bedrock, Anthropic, OpenAI, and every other connected provider, so finance sees one honest AI total instead of a line item buried in a cloud bill.

Read-only, cost-scoped credentials

CostMon connects with a service account scoped to bigquery.dataViewer, jobUser, and monitoring.viewer only. It's read-only and never touches your Vertex AI models, prompts, or infrastructure.

  • Per-tokeninput and output priced separately, and separately by model tier
  • ~50%typical Batch API discount vs. on-demand for the same model
  • Dailysync cadence, not monthly

Getting started

Connect, sync, see — for Vertex AI

  1. 01

    Connect a read-only GCP service account

    Grant a service account bigquery.dataViewer and jobUser on your billing export dataset, plus monitoring.viewer for token metrics. CostMon never requests access to your Vertex AI models or infrastructure.

  2. 02

    CostMon joins billing export and token usage

    Vertex AI rows are filtered out of your GCP billing export and joined to Cloud Monitoring's per-model input and output token counts on the same day, reconciled into the same normalized model as every other provider you connect.

  3. 03

    See the line before the invoice does

    Open a daily, per-model view of Vertex AI spend and catch a missed Batch discount or an uncached repeated-context workload weeks before it lands on a bill.

FAQ

Common questions about the Vertex AI connector

How does CostMon connect to Vertex AI?

Through the same read-only service account as the GCP connector (bigquery.dataViewer and jobUser on your billing export dataset), plus monitoring.viewer scoped to Cloud Monitoring's token-usage metrics. CostMon never requests access to your Vertex AI models or prompts.

Is Vertex AI spend separate from the rest of my GCP bill?

No. It's the same GCP Billing Export everything else reads from. CostMon filters it to Vertex AI / generative AI services and joins it with per-model token metrics, so it's broken out without needing a separate connector or invoice.

Can CostMon tell me if I should be using the Batch API?

CostMon surfaces per-model daily spend and volume, which is the signal you need to spot a high-volume, non-realtime workload running at standard online rates. It doesn't switch the workload to Batch for you, but it makes the opportunity visible.

Does CostMon track context caching savings?

Cached and standard input tokens both flow through Cloud Monitoring's token metrics, so a workload that isn't benefiting from context caching shows up as a higher effective input-token cost per call relative to similar workloads that are.

What does the Vertex AI connector cost?

A flat monthly rate per plan, not a percentage of your Vertex AI spend or the savings CostMon helps you find. Free includes 2 connectors; Professional and Enterprise include unlimited connectors.

Connectors

Cloud, AI, and SaaS: all in one place

AWS, GCP, and Azure for cloud; Anthropic, OpenAI, Amazon Bedrock, Vertex AI / Gemini, and Helicone for AI and LLM spend; Snowflake and Databricks for the data cloud; Datadog, GitHub, and Vercel for the rest of the stack. Add CSV import for any tool without a native connector. That's 14 connectors, all normalized into the same unified view.

Connect a provider. See one number.

Connect your first provider and CostMon normalizes it alongside everything else you run. Your team gets one number everyone can check.

Esc