Education · The LLM field guide
Meet the models.
Six frontier labs, two open-weight movements, and a couple of front doors most people use without knowing what's behind them. Here's each major large language model: where it came from, where it fits, and the issues around it that actually matter — including the standoff that reshaped who runs AI inside the U.S. government.
The roster
Who's who in the model wars.
ChatGPT / GPT
The company that made AI a household verb. Founded in 2015 as a research lab, OpenAI spent years publishing quietly before ChatGPT's launch in November 2022 became the fastest consumer-product adoption in history to that point. Its GPT models remain the default — the assistant most consumers try first, and the ecosystem with the broadest reach across apps, plugins, and integrations.
Where it fitsThe general-purpose default. Strongest all-around consumer experience, enormous third-party ecosystem, and — as of 2026 — the deepest footprint inside the U.S. federal government, with ChatGPT deployed to roughly three million defense personnel through the Pentagon's GenAI.mil platform and offered to civilian agencies for $1 a year through a GSA partnership.
OpenAI's history is punctuated by governance drama (the board fired and rehired CEO Sam Altman inside one week in November 2023) and copyright litigation from publishers over training data. Its February 2026 Pentagon agreement — signed within a day of Anthropic's federal ban — drew "delete ChatGPT" campaigns and internal criticism; Altman himself conceded the deal was "definitely rushed" and that "the optics don't look good." The company says the contract bars domestic mass surveillance and requires human control over uses of force.
Timeline
- 2015Founded as a nonprofit research lab
- 2020GPT-3 opens via API — developers get LLM access
- NOV 2022ChatGPT launches; ~100M users in ~2 months
- 2023GPT-4; Altman fired and reinstated in five days
- 2024GPT-4o goes multimodal — voice, vision, text
- 2025GPT-5 ships; agent features mature
- FEB 2026Pentagon classified-network deal; GenAI.mil rollout
- JUL 2026GPT-5.6 family (Sol / Terra / Luna)
Claude
Founded in 2021 by former OpenAI researchers — including siblings Dario and Daniela Amodei — Anthropic built its identity on AI safety, training Claude with a technique it calls constitutional AI. Claude earned its reputation in enterprise work: long documents, careful writing, and coding, where its models have traded benchmark leads with OpenAI's for two years.
Where it fitsThe enterprise and builder favorite — the model behind many professional coding tools, and a common choice for businesses that care about how their AI vendor governs itself. (Full disclosure of the obvious: it's also one of the model families we build with at Building Futures AI.)
This is the story that redrew the government AI map. Anthropic held a Pentagon contract worth up to $200M with two contractual red lines: its models could not be used for mass surveillance of Americans or for fully autonomous weapons. In late February 2026, after months of negotiation, the Pentagon demanded unrestricted "lawful use"; Anthropic's CEO said the company "cannot in good conscience accede." Within a day, the Defense Secretary designated Anthropic a supply-chain risk to national security — the first U.S. company ever given a label typically reserved for foreign adversaries' firms — and the President ordered all federal agencies to stop using its technology.
Anthropic sued, arguing the designation was retaliation that violates the First Amendment and due process; researchers from OpenAI and Google filed support for its safety position, and civil-liberties groups filed briefs on its side. The government argues a vendor can't dictate operational limits to the military. The litigation is ongoing — and the vacancy was filled almost immediately: OpenAI announced its own Pentagon deal the next day, and Google's Gemini was already on the military's platform. Whatever your politics, the business lesson is the same: model access can change overnight, for reasons that have nothing to do with technology.
Timeline
- 2021Founded by former OpenAI researchers
- MAR 2023Claude launches
- 2024Claude 3 family; enterprise adoption accelerates
- 2025Claude 4 generation; coding-tool dominance; Pentagon contract
- FEB 2026Refuses to drop safeguards; federal ban; lawsuit filed
- 2026Claude 5 generation ships amid the litigation
Gemini
Google invented the transformer architecture every modern LLM is built on (2017), then watched ChatGPT beat it to market. Its answer — Bard, then Gemini — started shaky and matured fast, and Google's real advantage was never the chatbot: it's distribution. Gemini is wired into Search, Maps, Android, Workspace, and Google Business Profiles.
Where it fitsThe model a local business can least afford to ignore. Gemini's answers draw on Google's index and Business Profile data — meaning your GBP hygiene now feeds an AI's recommendations directly. It was also the first frontier model on the Pentagon's GenAI.mil platform in late 2025, making Google a quiet giant in government AI.
Gemini's ancestor Bard flubbed a fact in its own launch demo (briefly denting Alphabet's market value), and 2024's AI Overviews launch produced infamous errors before settling down. The deeper, unresolved tension: Google's AI answers reduce clicks to the websites the answers are built from — a structural fight with publishers and small businesses that's still playing out.
Timeline
- 2017Google researchers publish the transformer paper
- FEB 2023Bard launches in ChatGPT's shadow
- DEC 2023Rebuilt and renamed Gemini
- MAY 2024AI Overviews land in Google Search
- LATE 2025Gemini for Government — first on GenAI.mil
- 2026Gemini 3.5 generation; deep Workspace integration
Grok
Elon Musk's 2023 entry, built at startling speed and wired directly into X (Twitter). Grok's signature is real-time social data — it knows what's happening right now in a way models trained on static snapshots can't — plus a deliberately irreverent voice.
Where it fitsLive events, social sentiment, and breaking context; recent versions are also genuinely competitive on coding. For most local businesses it's a secondary channel — but its answers still recommend companies, so it belongs in any visibility tracking.
Grok's loose guardrails have produced repeated public incidents — most infamously a July 2025 episode where the model generated antisemitic content and praised Hitler before xAI intervened. The pattern raises the question every business should ask of every AI vendor: who's accountable when the model speaks?
Timeline
- 2023xAI founded; Grok 1 ships within months
- 2024Grok 2–3; image generation; X integration deepens
- JUL 2025Content scandal forces guardrail overhaul
- 2025Grok 4 — benchmark-competitive
- 2026Grok 4.5; aggressive API pricing
Llama
Meta took the opposite bet: give the weights away. Llama models can be downloaded and run on your own hardware — no API, no per-token fees, full data control. Llama 4 (2025) pushed context windows to extremes and seeded thousands of fine-tuned variants across the industry.
Where it fitsSelf-hosting, privacy-sensitive work, and cost control at scale — the backbone of the open-weight ecosystem.
Open weights mean anyone can strip safety training out — the industry's longest-running governance debate. And "open" comes with an asterisk: Meta's license has commercial limits, and you carry the ops burden.
DeepSeek V / R
The January 2025 shock: a Chinese lab released R1, a reasoning model near frontier quality, trained at a fraction of reported U.S. costs — and briefly wiped roughly $600B off Nvidia's market value in a single day as markets repriced what AI has to cost. Its 2026 models keep resetting the price floor.
Where it fitsHigh-volume, cost-sensitive workloads via API or self-hosting — the reason frontier API prices keep falling for everyone.
Data governance: it's a Chinese company subject to Chinese law, and multiple governments have restricted official use of its hosted apps. Open-weight self-hosting sidesteps some concerns — but the diligence burden is yours.
Copilot & Perplexity
Two names your customers use that aren't really model makers. Microsoft Copilot wraps OpenAI's models (plus Microsoft's own) inside Word, Excel, Outlook, and Windows — for millions of office workers, Copilot simply is AI. Perplexity is an answer engine that searches the live web, runs your question through frontier models, and cites its sources — which makes it the most transparent window into how AI composes a recommendation, and a favorite tool in our own visibility audits.
Why they matter to youThese are where model choice disappears from the customer's view. Someone asking Copilot or Perplexity "best restoration company near me" neither knows nor cares which model answers — but the answer still names somebody. That's why visibility work has to span the whole roster, not one chatbot.
What it means for you
Multi-model is the only sane strategy.
Three takeaways from the roster. First: no single model "won" — customers are spread across all of them, so your visibility has to be measured across all of them. Second: the federal fight proved that access to a model can change overnight for non-technical reasons; systems built with a single hard-wired vendor carry real platform risk, which is why we architect client systems to swap models. Third: the differences between models are exactly why tracking matters — the platform wired into Google's index and the one searching the live web can give different answers about your business on the same day.