The Toll Matrix

A two-axis grid of an independent AI and technology research corpus: seven coverage sub-sectors crossed against seven altitude layers of one compute stack, from the physical substrate that creates compute to the applications that spend it. It documents structural coverage breadth; it does not make recommendations.

Point-in-time rendering, July 23, 2026. Underlying framework authored June 2026; classification pass July 2026.

Rows are sub-sectors as the source research groups them, 57 company write-ups in total. Columns are altitudes: where in the flow of compute an occupant actually collects. The two axes are deliberately independent, and the tension between them is the point: a write-up filed under enterprise software can sit at the bottom of the physical substrate, and one sub-sector's governing economics refuse to map onto the axis at all, which the grid says honestly instead of forcing a fit. Cells record the mechanism at each intersection, and name the notable public occupants of each layer as any observer of the industry would name them: the roster is derived from public knowledge of who operates where in the stack, is independent of the research corpus the counts describe, and implies no mapping between named companies and the write-up counts. Every company description on this page was written from public knowledge. Click any marked cell for the occupants and the methodology that placed them.

Physical substratecreating compute
Compute landlordrenting capacity
Compute executionaddressing silicon
Dataquality and gravity
Model weightsthe thin layer
Routing and controlconnectivity
Applicationthe end task
Semiconductors12 write-ups
Optical interconnect8 write-ups
Physical infrastructure12 write-ups
Hyperscalers3 write-ups
AI labs1 write-up
Enterprise software17 write-ups
Industrial AI4 write-ups

mechanism with named public occupants   the hinge (contested)   badges: P perimeter · C content gate · M meter and governance   click any marked cell

The hinge: two standards, opposite directions

The compute execution layer and the routing layer run the same play: author a standard, let everyone build on it, and let the accumulated gravity do the selling. Arm runs the version one floor down, licensing processor architecture into most of the world's smartphones and roughly a third of all processor chips. But the two software chokepoints run the play in opposite directions.

CUDA, NVIDIA's programming platform, locks compute in. In practice the dominant path to the GPU for AI workloads runs through the programming model, so the switching cost points down toward the silicon and the toll is collected on every training and inference cycle. The Model Context Protocol, authored by Anthropic and donated in December 2025 to the Agentic AI Foundation (a directed fund under the Linux Foundation), lets data out: it collapses M-by-N integration work to M-plus-N, commoditizes the connectors, and pushes competition up to model quality. One toll is a wall; the other is a door.

Open compiler paths (PyTorch 2.x, Triton, AMD's ROCm) challenge the training-side lock, while deployment runtimes (TensorRT and version-pinned inference engines) deepen the serving-side lock at exactly the moment a product ships and latency starts to matter. The grid records both dynamics. Which prevails is a question it deliberately does not answer.

Two fault lines the grid records but does not resolve

Fault line 1: metered gate versus flat license

The content gate splits on billing form. A per-crawl meter collects in proportion to use (TollBit and Cloudflare's pay-per-crawl program are the two most visible implementations between publishers and AI crawlers); a flat license collects once regardless of use. The two forms allocate usage risk differently between buyer and seller; the grid records the split as structural and stops there.

Fault line 2: the execution-layer contest

Challenge at training, deepening lock at serving, as described in the hinge note above. Both readings are carried simultaneously; the column is carried as contested (see the Altitude 3 key) rather than silently resolved in either direction.

Column key

Altitude 1Physical substrateThe tax on creating compute

A chain of narrow markets, most with one, two, or three credible suppliers: chip-design software, architecture licensing, lithography and wafer-fab equipment, foundry and advanced packaging, high-bandwidth memory, and the power, cooling, and optical assembly that make rack density physically possible. Optimization does not shrink this toll: compressing the per-token cost of inference expands the workloads run on top of it (the Jevons pattern), which increases load on downstream memory and storage. A published NVIDIA research method compressing attention KV caches roughly 20x with model weights unchanged (KVTC, arXiv:2511.01815, accepted to ICLR 2026) is the current reference case for how fast that compression is moving.

Mechanism

Every chip that exists paid this column first; the toll is collected on creation, not on use.

Binding constraint

Physics and installed capacity. You cannot code around lithography, thermodynamics, or packaging throughput.

Altitude 2Compute landlordThe tax on renting standing capacity

Rent on the metal: the building, the power contract, the racks, sold as raw accelerator capacity or as capacity plus an inference platform. The landlord's economics are real-estate economics priced in compute, which is why the growth ceiling is grid build velocity rather than AI demand.

Mechanism

Instance-hour billing on standing capacity; the toll survives model churn because the tenant changes and the building does not.

Binding constraint

Energization speed: how fast contracted power becomes connected power.

Altitude 3Compute executionThe tax on addressing the silicon; the hinge

The seam between the hardware and software halves of the stack. Whoever owns the programming model that addresses the accelerator collects on every FLOP that crosses it, which makes this layer simultaneously the top of the silicon story and the floor of the software story: a multi-layer resident, not a filing ambiguity. Full treatment in the hinge note above.

Mechanism

The programming model as gravity: everything built on it deepens it.

Binding constraint

Contested by design: challenged at training, deepening at serving.

Altitude 4DataThe tax on quality, not speed

Extraction and preparation of raw corporate documents into AI-ready form, streaming and operational stores, and retrieval optimization between the store and the model. The toll is real because bad data induces failure: enterprises must structure their pipelines before agents can be trusted with them. It is not uniformly a clean toll; pricing models built for pre-agent query volumes can be squeezed as agentic workloads change consumption patterns.

Mechanism

Data gravity plus correctness: traffic routes through whoever makes the corpus usable and keeps it that way.

Binding constraint

Correctness at scale; the content gate wraps this column hardest.

Altitude 5Model weightsThe thinnest toll in the building

The one stretch of the stack that everyone mistakes for the destination. Substitution is one API pointer away, open-weight releases keep compressing per-token pricing toward marginal cost, and no constraint at this altitude is durable by itself. Occupants of this layer often also appear one column to the right, in routing and standards, where the structural mechanics differ; the grid records both placements.

Mechanism

Per-token pricing on frontier quality, eroded continuously by open-weight parity.

Binding constraint

None durable; that is the finding, not a gap in it.

Altitude 6Routing and controlThe tax on connectivity and orchestration

Protocols, gateways, sandboxes, and the payment rails agents use to pay for external services. Gateways that store every model's API keys and cache upstream calls have quietly become critical infrastructure; a March 2026 supply-chain compromise at LiteLLM, a widely deployed open-source gateway (trojanized packages published through a hijacked maintainer account, disclosed by the project and documented by multiple security firms), promoted the category from utility router to security perimeter in a single incident. One recurring structural pattern at this altitude is authoring a standard rather than owning a product.

Mechanism

Standards gravity and chokepoint custody: connectivity must route through something, and the something accumulates.

Binding constraint

Trust and standardization; the meter and governance wrap concentrates here and in the application column.

Altitude 7ApplicationThe tax on the end task

The altitude where the toll is hardest to keep. Agents erode seat pricing, so survival runs through owning the system of record beneath the agent and re-pricing from seats to hybrid metering that passes compute cost through rather than absorbing it. Most application software does not carry a durable toll; the survivors own the record the agent must write back to.

Mechanism

Metered work on the end task, collected only where the system of record underneath is owned.

Binding constraint

Seat erosion; the pass-through test decides who keeps pricing.

WrapsThe three cross-cutting tollsCollected at every altitude, shown as badges

Three tolls wrap the stack rather than sitting at one altitude, so the grid shows them as per-cell badges instead of columns. Perimeter (P): security as a non-discretionary risk tax that grows with machine identities, because every agent is an endpoint to authenticate. Content gate (C): metered or licensed access to legally cleared corpus, split down the middle by billing form (fault line 1 above). Meter and governance (M): observing and governing agentic flow, because you cannot manage what you cannot trace.