The AMICE Stack and the Far Side of Model Mountain
Applications, Models, Infrastructure, Chips, Energy. Most AI users rarely see below Applications, even though the lower four layers decide what the top one can do.
Applications, Models, Infrastructure, Chips, Energy. Most AI users rarely see below Applications, even though the lower four layers decide what the top one can do.
When people argue about AI, they usually compare apps and models. The money, physical limits, and geopolitics continue below them through infrastructure, chips, and energy: the five layers I’m calling AMICE.
Most everyday AI arguments are about the app someone opened that morning and the model behind it. ChatGPT or Claude. Which one writes the better email. I have those arguments too.
AMICE is my way of zooming out: Applications, Models, Infrastructure, Chips, Energy. I did not invent the underlying stack. Jensen Huang uses the same five layers, starting from the other end:
“AI is a five-layer cake. Energy, chips, infrastructure, models, and applications.”
— Jensen Huang, AI Is a 5-Layer Cake, NVIDIA
The useful bit is not the acronym. It is remembering that a prompt reaches much farther than the chat window.
AMICE
The stack moves in two directions
Applications are the chat apps, coding agents, and AI features inside other software. This is the only layer where most people see a button, a price, and a product name.
Models are GPT, Claude, Gemini, and the open-weight families. Frontier training costs hundreds of millions of dollars, so there are far fewer model builders than application developers.
Infrastructure is the data centers: land, buildings, cooling, networking, and the software that makes a hundred thousand GPUs behave like one machine. A rack of expensive chips is not useful if it cannot be powered, cooled, connected, and scheduled.
Chips are NVIDIA's GPUs, Google's TPUs, and the fabrication and packaging capacity behind them. Most leading accelerators eventually pass through TSMC, which is why Taiwan keeps appearing in arguments about AI policy. Dario Amodei spent an entire essay arguing about export controls because access to chips changes who can train at the frontier.
Energy is generation, transmission, substations, and the contracts that connect a data center to all three. The IEA has a whole report on energy and AI because data-center demand now belongs in grid plans.
Demand moves down the stack. A popular app needs more inference, which needs more racks, accelerators, and electricity. Shortages move back up. If a site cannot get enough power, the cluster cannot expand, the model provider has less capacity, and the app starts rationing requests. Sam Altman compresses this into the claim that compute will become the currency of the future. Satya Nadella uses a more practical measure:
“Tokens per watt per dollar.”
— Satya Nadella, his shorthand for where energy, compute, and cost meet
It connects the thing a user receives to the electricity and money spent producing it.
Trace one prompt
Step 1 of 5 · Applications
A prompt starts with a person.
The visible part is the app: the chat window, coding agent, or feature where demand first appears.
The model picker belongs to the app
Model menus, subscriptions, and features are the choices in front of you. A new application can ship without buying land, building a fab, or negotiating a power contract.
Even a model picker is application UI. The product team chose which models appear, what context they receive, how much capacity each one gets, and what the subscription costs. I still think the best AI model is the one you actually use. I just do not want to confuse choosing from a menu with running the kitchen.
I needed a picture for that separation, which is how I ended up with Model Mountain. Applications and model choices crowd the near side. Beyond the ridge are data centers, chip fabs, interconnects, substations, and power plants. I will probably spend my career on the near side. Knowing what is across the ridge still helps me explain a rate limit, a price increase, or why a promised feature is taking longer than expected.
In 2022 and 2023, “that’s just a ChatGPT wrapper” became the standard way to dismiss an AI startup. Sometimes it was accurate. Plenty of products were a prompt, an API call, and a new landing page. TechCrunch was warning founders that integration alone was not a business. If OpenAI could reproduce the whole product in a release note, there was not much protection underneath it.
The phrase stopped being useful when “wrapper” came to mean any application built on somebody else’s model. By that definition Cursor is a wrapper. So is nearly every useful piece of AI software I use. Distribution, workflow depth, customer data, support, and all the unglamorous software around the API call still matter. By 2025, Forbes was arguing only fools use “ChatGPT wrapper” as an insult. I would not go that far, but the old insult no longer tells me whether a product is any good.
I made the stack because “AI news” is too broad to be useful. A model launch, a chip export rule, and a power contract change different parts of the system on different timelines.
A model release starts at M, but I usually experience it only when an application ships it. An export rule starts at C and may change which models can be trained later. In 2024, Constellation signed a 20-year power agreement with Microsoft to support restarting Three Mile Island Unit 1. Microsoft will buy the 835-megawatt plant’s output to match electricity used by its data centers in the PJM grid region. The buyer is a software company, but that headline belongs under Energy. Putting all three examples in one “AI” bucket hides what actually changed.
If someone says AI progress will stall, I want to know what they expect to run short: users, model research, data-center capacity, accelerators, or megawatts. Without that, I do not know what evidence would prove the forecast right or wrong.
A small application can begin with a few engineers and an API account. Siting a data center, adding fab capacity, or building new generation takes more capital, permits, and time. The people I want to hear from are often working between those clocks: a product engineer who understands inference budgets, or a data-center planner who knows which workloads developers are actually trying to run.
I will still spend almost all of my AI time in applications. But when one gets slower, more expensive, or more heavily rate-limited, I have better questions to ask. Is the provider short on accelerator time, rack space, or power? AMICE does not answer the question, but at least it tells me where to look.
Want to see what sits underneath a product you use? Open the interactive AMICE Explorer, or choose any layer in the diagrams above.