The Latent
EN▾
EnglishEspañol中文PortuguêsFrançaisالعربية日本語한국어
Sign Up
NEWSBENCHMARKSDATALEARNNEWSLETTERPARTNER WITH US →
Data/Capability

Capability

Maximum API Model Context Window Over Time

Record-high provider-stated input/context capacity of broadly available hosted language-model APIs since 2023. Larger windows expand how much text, code or document material can be supplied per request, shaping application design and token costs; they do not measure output length or effective retrieval quality.

Maximum API Model Context Window Over Time

The documented record API context window was 10,000,000 tokens as of 2026-08-27.100M10M{"f":[800,420,56,16],"s":[["Record maximum context","#9333ea"]],"p":[["2026-08-27T00:00:00.000Z","Aug '26",56,[[0,"10,000,000 tokens",380,null]],null]]}Record maximum context: 10M on Aug '26
SOURCE: Official provider announcements and documentation, with Wayback corroboration
RANGEALLYTD12M3M1M
Key takeaway

The documented commercial-API record rose from Anthropic's 100K-token window in May 2023 to Alibaba Cloud's 10M-token Qwen-Long service by May 2024, expanding the maximum material an application could submit in one model request. Qwen-Long requires file-ID submission above 1M tokens, and 10M is a provider-stated capacity—not an independent measure of retrieval quality.

The documented record API context window was 10,000,000 tokens as of 2026-08-27.

Pro API coming soon

Methodology

Scope begins January 1, 2023 and is reviewed through August 27, 2026. A milestone is added only when a provider-published limit exceeds every earlier qualifying limit. One equal-value carry-forward observation at the review cutoff extends the final plateau and is explicitly labeled as an audit point, not a new record.

Qualifying models must be callable through a first-party hosted commercial API, not merely described in a paper, benchmark, downloadable checkpoint, product demo, or research preview. A rollout to existing API customers qualifies. A provider-labeled preview qualifies only when the provider explicitly opened it to all paying API developers; invite-only, waitlist, private-preview, and bespoke research access do not.

The timestamp is the provider's dated announcement or the earliest retained official evidence of broad API availability. It is not backfilled from a later news report. For Qwen-Long, May 21, 2024 is anchored to Alibaba Cloud's official notice that lists Qwen-Long on Model Studio and says the new API pricing became effective that day; the linked official model page supplies the 10M input/context specification.

Values reproduce the provider's stated decimal token capacities: 100K is stored as 100,000, 128K as 128,000, 200K as 200,000, and 10M as 10,000,000. Binary token counts are not inferred when the announcement uses rounded decimal notation.

The plotted measure is the maximum context or input capacity accepted for one model request. Output limits are not added to the plotted value. GPT-4 Turbo's documented maximum output was 4,096 tokens; current Qwen-Long model documentation lists 8,192 maximum output tokens. Historical Anthropic announcements described output separately but did not provide a stable exact output ceiling on the cited launch pages.

Qwen-Long's 10M path is operationally different from a single raw HTTP body: Alibaba documents direct HTTP text submission up to 1M tokens and requires larger documents to be uploaded and referenced by file ID. The provider says those referenced tokens count as request input and are used for model inference, so the published 10M maximum input/context is retained with this caveat.

Context capacity is not effective context. Needle-in-a-haystack recall, long-context benchmark scores, hallucination reductions, and claims that a model can find or use information across the full window are performance evidence and are never converted into token capacity.

The chart is a record frontier, not a catalog. Non-record releases are intentionally absent except for the review-cutoff carry-forward point. Lines connect discrete documentary milestones and do not imply that capacity changed gradually between dates.

Frequently asked questions

Does context window mean maximum input or maximum output?

No single convention is universal, so this chart uses the provider-stated request context or maximum input capacity and keeps output ceilings separate. A model can accept millions of input tokens while producing only thousands of output tokens.

Why does maximum context capacity matter to the AI industry?

A larger input window can let applications submit more documents, code, conversation history or other source material in one request, affecting retrieval architecture, workflow design and token-based costs. Capacity alone does not show that a model can accurately find, reason over or remember every part of that input.

How should a rise in the record be interpreted?

It means a broadly available first-party hosted API documented a higher maximum request context or input limit than earlier qualifying services. It marks a change in the commercial API capacity frontier, not a proportional improvement in answer quality, latency, affordability or usable long-context reasoning.

Why is GPT-4 Turbo included even though OpenAI called it a preview?

OpenAI explicitly made gpt-4-1106-preview available to all paying API developers on November 6, 2023. The inclusion rule accepts provider-labeled previews with broad commercial API access, while excluding waitlists and invite-only previews.

Why are Gemini's 1M and 2M windows not points?

Google's 1M private preview in February 2024 is excluded by the access rule. By the time Gemini 1.5 Pro's 1M model and later 2M window reached general availability in May and June 2024, official evidence already placed Qwen-Long at 10M on May 21, so neither set a new record.

Is Qwen-Long's 10M figure a normal prompt?

It is an official API input/context limit, but inputs above 1M use Alibaba Cloud's upload-and-file-ID workflow because of request-body limits. The files are parsed before inference; Alibaba says their tokens are counted as model input on every call. This chart does not claim that retrieval remains equally accurate at every position.

Why are Magic's 100M model and Llama 4 Scout absent?

Magic's LTM-2-mini announcement was a research update without a generally available hosted API. Meta's Llama 4 Scout is a downloadable 10M-context model and therefore does not exceed the existing 10M record; first-party broad API evidence for the full 10M limit was also not used.

Why is MiniMax-Text-01 absent?

MiniMax announced a 4M inference context and commercial API pricing on January 15, 2025, but 4M was below the 10M record already documented for Qwen-Long.

Has the 10M record been independently tested?

No. The series records official API contracts and announcements, not an independent benchmark. Effective retrieval and reasoning can degrade well before a nominal window is full.

How are milestones sourced and selected?

Each increase requires first-party capacity documentation plus dated official evidence of broad hosted-API availability. The chart excludes papers, downloadable checkpoints, demos, invite-only access and later releases that merely tie or fall below the standing record. Dates are not backfilled from secondary reports.

Related charts

  • AI Benchmark Saturation Curves
  • Largest Documented Training Compute by Release Year: China vs. U.S.
  • Frontier Model Training Compute

The Latent

AI industry news. A sister publication to The Block.

Editorial

  • Standards
  • Corrections
  • Commercial policy
  • Contact

Company

  • LEARN
  • Data
  • Benchmarks
  • Score
  • Methodology
  • About
  • Team
  • Privacy Policy
  • Terms of Service
  • Security
  • The Block
  • Add The Latent as a preferred source