Ghost Technology
Insights
FinOps · June 2026

The AI bill became a black box. Token FinOps is how you open it.

PointFive just raised $60 million and shipped a platform built for the one part of the cloud bill that reserved instances cannot touch. Here is what enterprise buyers should actually take from it.

6 min read
In partnership withPointFive

Every enterprise running generative AI in production has quietly created a new problem on its cloud bill. Azure OpenAI, AWS Bedrock, and Google Vertex AI tend to report spend as a single aggregated line. You see the total. You cannot see which model, which deployment, or which team produced it. As AI moves from pilot to production, that gap stops being an accounting inconvenience and becomes a budgeting risk.

// The playbook that broke

Reserved capacity logic does not survive contact with tokens

For more than a decade, cloud cost management ran on a stable assumption. A workload lived on an instance, the instance belonged to a team, and you bought down the rate with reserved instances or savings plans. Token based billing breaks every link in that chain. There is no instance to reserve. Cost depends on input tokens, output tokens, cached tokens, model choice, and provisioned throughput, and a small change to a prompt or a model can swing spend by a large amount.

The visibility problem is worse than the pricing problem. Configuration happens at the deployment level, while cost aggregates at the account level. The engineers who can actually fix the spend are the ones who cannot see it. The playbook that works cleanly for compute and storage does not map onto inference, which is exactly why a dedicated discipline is forming around it.

Reserved era

Rate optimization

You committed to instances and bought the rate down. The unit was the server, it belonged to one team, and a savings plan made it cheaper for a year. Predictable, and increasingly beside the point for anything running inference.

Token era

Unit economics

There is no instance to reserve. Cost moves with input tokens, output tokens, cached tokens, and model choice. The only way to manage it is to measure it, per model, per deployment, and per team, then act on what you find.

// What PointFive shipped

A detection engine that ships its own fixes

PointFive is one of the clearer signals that this category is maturing into real infrastructure. On June 8 at FinOps X in San Diego, the company launched what it calls the AI Efficiency OS along with a companion product, TokenShift, in the same week it announced a $60 million Series B led by Accel. The positioning is deliberately broad: efficiency that follows waste from the cloud all the way to the coding agent.

Underneath sits a detection engine the company calls DeepWaste, with more than 500 validated detections spanning cloud infrastructure, data platforms, GPU capacity, and managed LLM services. For AI specifically it maps the full cost surface, from token level economics on Bedrock, Azure OpenAI, and Vertex AI to GPU rightsizing and provisioned throughput tuning. It flags reserved capacity sitting idle, surfaces model migrations where a newer model delivers better unit economics, and breaks cost down per input token, output token, and cached token. It connects read only, with no data plane access.

The part that separates it from a dashboard is what happens next. Rather than handing engineers a list of recommendations, PointFive drafts the pull request, opens the ticket, routes the owner in Slack, and verifies the fix landed in production before booking the savings back to the FinOps ledger. TokenShift extends the same accounting into the IDE, so the tokens that coding agents now burn are governed where they are actually spent.

How it works

From read-only connection to verified savings

01

Read-only connect

PointFive authorizes against your environment with no data plane access. Findings begin once the integration is live.

AWSAzureGCPKubernetesSnowflakeDatabricks
02

DeepWaste detection

The engine reads configuration, telemetry, and code, then surfaces the AI waste that generic cost tools miss. More than 500 validated detections.

Token-level economicsPTU and GPU rightsizingIdle capacityModel migration
03

Agentic remediation

Every finding carries an owner, a risk profile, and a savings estimate, then moves itself toward the fix instead of waiting in a dashboard.

Draft pull requestOpen ticketPage owner in SlackVerify in production
04

Savings booked to the ledger

Realized savings are written back to the FinOps ledger. The loop runs continuously as models, prompts, and capacity change.

PointFive detect and remediate loop. Based on PointFive platform documentation, June 2026.

A platform that finds savings you have no engineering capacity to ship is a report, not a return.
// The buy versus build call

This is a buy, not a build

Our advice on this is consistent. Token FinOps is a real discipline, and for almost every enterprise it is a buy, not a build. The attribution problem is genuinely hard, the billing mechanics shift every quarter, and the engineering hours it takes to reconstruct this internally are almost never worth it. Teams that try usually end up with a brittle dashboard nobody trusts and a roadmap item that never ships.

The fair question is not whether the problem is real. It is whether you already have enough inference and GPU spend, and the engineering capacity to act on what a platform surfaces, to justify a dedicated tool right now. That is the conversation we would rather have than a sales pitch. When PointFive is the right fit, we will tell you. When it is still early for your environment, we will tell you that too.

// Where Ghost sits

Why we brought PointFive to the table

We spend our time on this category because AI cost is moving from an engineering footnote to a board-level line faster than most teams have priced in. We brought PointFive to our customers because the token attribution problem is real and the platform answers it well, and telling you what is worth buying and what is not worth building is the whole reason we exist. Having a point of view on where the spend is going, and being willing to stake it, is what we think a partner is actually for.

// Resources

Go deeper.

Blog · PointFive

The tokenomics frontier

Why per-token unit economics, not aggregate spend, are the real unit of management for AI workloads.

Read more →
Blog · PointFive

GenAI unit economics across every cloud

Cross-cloud comparison frameworks and attribution models for AI cost governance on Bedrock, Azure OpenAI, and Vertex AI.

Read more →
Announcement · PointFive

The AI Efficiency OS and TokenShift

PointFive's June launch: an AI-native platform plus a companion product that governs token usage wherever coding agents run.

Read more →

Talk to Ghost.

If AI cost is turning into a board-level line, we can map which platform fits your environment and which marketplace retires the most committed spend.