The AMICE stack and the far side of Model Mountain
I use AMICE to connect AI applications and models to the infrastructure, chips, and energy underneath them. Those dependencies help explain costs, delays, and limits.
I use AMICE to connect AI applications and models to the infrastructure, chips, and energy underneath them. Those dependencies help explain costs, delays, and limits.
AI arguments usually compare apps and models. I'm calling the wider system AMICE because infrastructure, chips, and energy also determine what those products can do.
Most everyday AI arguments are about the app someone opened that morning and the model behind it. ChatGPT or Claude? Which one writes the better email? I have those arguments too.
AMICE stands for Applications, Models, Infrastructure, Chips, and Energy. I didn't invent the underlying stack. In his March 2026 essay AI Is a 5-Layer Cake, Jensen Huang describes the same layers in the opposite order, starting with energy. It is a useful industry framing from the head of a company that sells much of the hardware underneath it.
A prompt reaches much farther than the chat window.
AMICE
The stack moves in two directions
Applications are the chat apps, coding agents, and AI features inside other software. This is the only layer where most people see a button, a price, and a product name.
Models are GPT, Claude, Gemini, and the open-weight families. Training at the frontier can require enormous capital. Epoch AI's cost study estimates hardware, energy, and staff costs separately, which is a better starting point than treating the price of one training run as the entire cost of making a model.
Infrastructure is the data centers: land, buildings, cooling, networking, and the software that makes a hundred thousand GPUs behave like one machine. A rack of expensive chips needs power, cooling, connections, and scheduling to be useful.
Chips are NVIDIA's GPUs, Google's TPUs, and the fabrication and packaging capacity behind them. Most leading accelerators eventually pass through TSMC, which is why Taiwan keeps appearing in arguments about AI policy. Dario Amodei spent an entire essay arguing about export controls because access to chips changes who can train at the frontier.
Energy is generation, transmission, substations, and the contracts that connect a data center to all three. The IEA has a whole report on energy and AI because data-center demand now belongs in grid plans.
Demand moves down the stack. A popular app needs more inference, which needs more racks, accelerators, and electricity. Shortages move back up. If a site cannot get enough power, the cluster cannot expand, the model provider has less capacity, and the app starts rationing requests. Sam Altman compresses this into the claim that compute will become the currency of the future. Satya Nadella uses a more practical measure:
"Tokens per watt per dollar."
Satya Nadella, his shorthand for where energy, compute, and cost meet
It connects the thing a user receives to the electricity and money spent producing it.
Trace one prompt
Step 1 of 5 · Applications
A prompt starts with a person.
The visible part is the app: the chat window, coding agent, or feature where demand first appears.
The diagram is a way to organize dependencies, not a literal network trace. A model can run on a phone. A cloud can buy accelerators from several suppliers. One application can call several models before returning an answer. Data, people, capital, and regulation cut across all five layers.
That distinction helps with a deceptively simple headline: a model lab "has" a million GPUs. Does it own them, rent access, or hold a contract for future capacity? Epoch AI's September 2026 Chip Users explorer separates compute use from ownership and reports estimates in H100-equivalents with uncertainty ranges. An equivalent is a comparison unit, not an inventory of physical H100s. I want the headline to distinguish ownership from access.
The model picker belongs to the app
Model menus, subscriptions, and features are the choices in front of you. A new application can ship without buying land, building a fab, or negotiating a power contract.
Even a model picker is application UI. The product team chose which models appear, what context they receive, how much capacity each one gets, and what the subscription costs. I still think the best AI model is the one you actually use. I just don't want to confuse choosing from a menu with running the kitchen.
I needed a picture for that separation, which is how I ended up with Model Mountain. Applications and model choices crowd the near side. Beyond the ridge are data centers, chip fabs, interconnects, substations, and power plants. I'll probably spend my career on the near side. Knowing what is across the ridge still helps me explain a rate limit, a price increase, or why a promised feature is taking longer than expected.
In 2022 and 2023, "that's just a ChatGPT wrapper" became the standard way to dismiss an AI startup. Sometimes it was accurate. Plenty of products were a prompt, an API call, and a new landing page. TechCrunch was warning founders that integration alone was not a business. If OpenAI could reproduce the whole product in a release note, there was not much protection underneath it.
The phrase stopped being useful when "wrapper" came to mean any application built on somebody else's model. By that definition Cursor is a wrapper. So is nearly every useful piece of AI software I use. Distribution, workflow depth, customer data, support, and all the unglamorous software around the API call still affect the product. By 2025, Forbes was arguing only fools use "ChatGPT wrapper" as an insult. I wouldn't go that far, but the old insult no longer tells me whether a product is any good.
I made the stack because "AI news" is too broad to be useful. A model launch, a chip export rule, and a power contract change different parts of the system on different timelines.
A model release starts at M, but I usually experience it only when an application ships it. An export rule starts at C and may change which models can be trained later. In 2024, Constellation signed a 20-year power agreement with Microsoft to support restarting Three Mile Island Unit 1. Microsoft will buy the 835-megawatt plant's output to match electricity used by its data centers in the PJM grid region. The buyer is a software company, but that headline belongs under Energy. Putting all three examples in one "AI" bucket hides what actually changed.
If someone says AI progress will stall, I want to know what they expect to run short: users, model research, data-center capacity, accelerators, or megawatts. Without that, I don't know what evidence would prove the forecast right or wrong.
A small application can begin with a few engineers and an API account. Siting a data center, adding fab capacity, or building new generation takes more capital, permits, and time. The people I want to hear from are often working between those clocks: a product engineer who understands inference budgets, or a data-center planner who knows which workloads developers are actually trying to run.
I'll still spend almost all of my AI time in applications. But when one gets slower, more expensive, or more heavily rate-limited, I have better questions to ask. Is the provider short on accelerator time, rack space, or power? AMICE does not answer the question, but at least it tells me where to look.
Want to see what sits underneath a product you use? Open the interactive AMICE Explorer, or choose any layer in the diagrams above.
These reading notes were checked on September 17, 2026. The interactive source ledgers retain their separately dated records.
NVIDIA: AI Is a 5-Layer Cake. The original industry framing. Read it as a chip supplier's argument about the system.
Epoch AI: the rising costs of training frontier models. A research cost model with explicit assumptions. Estimates and extrapolations are not audited invoices.
IEA: Key Questions on Energy and AI. The 2026 demand outlook and its supply constraints. Keep all-data-center totals separate from AI workloads.