Building frontier modelsthat scale efficiently

Subquadratic is a frontier AI research company building the most compute-, memory- and sample-efficient algorithms and models.

Built by researchers and engineers from
  • Meta
  • Google
  • Oxford
  • Cambridge
  • MIT
  • BYU
  • Amazon
  • DoorDash
  • LlamaIndex
  • TwelveLabs
  • Airbnb

Introducing SubQ
The best model for data-intensive workloads.

SubQ is the first model built for multi-million token reasoning, allowing enterprises to work across full repositories, financial filings, and contract archives for a fraction of the cost.

Built by researchers and engineers from
  • Meta
  • Google
  • Oxford
  • Cambridge
  • MIT
  • BYU
  • Amazon
  • DoorDash
  • LlamaIndex
  • TwelveLabs
  • Airbnb
The problem

Enterprises are spending billions to work around the architectural limitations of today's models.

  • Token Cost
  • Context Rot
  • Latency
  • Chunking
  • Embedding
  • Vector DB
  • Context Compression
  • Hallucinations
  • Eval Patches
  • Context Drift
  • Alignment Gaps
  • Token Cost
  • Context Rot
  • Latency
  • Chunking
  • Embedding
  • Vector DB
  • Context Compression
  • Hallucinations
  • Eval Patches
  • Context Drift
  • Alignment Gaps
  • Token Cost
  • Context Rot
  • Latency
  • Chunking
  • Embedding
  • Vector DB
  • Context Compression
  • Hallucinations
  • Eval Patches
  • Context Drift
  • Alignment Gaps
Use cases

The data-heavy workloads that become possible.

Whole Artifact Reasoning

Analyze entire datasets at once: an entire GitHub repo, all legal documentation, every vendor contract, or years of SEC filings — without chunking, compression, or context loss.

Long Horizon Agents

Agents that hold the full task in context and remember long sequences of events, collapsing brittle multi-agent orchestration into one coherent run.

Search & Retrieval

Surface material insights that you can trust from across an entire company's IP, code, or documents without risk of hallucination.

All your context.
Always available.

Reason across millions of tokens in one prompt: entire repos, whole artifacts, and long-running agent state, with context to spare.

Five years of SEC filings

A Fortune 500 10-Ks and 10-Qs

Python source code

The entire 3.13 standard library

Six months of React PRs

~1,050 pull requests against the React codebase

~ Approximate token counts.

The breakthrough

The most compute efficient model.

Today's transformer-based LLMs waste compute by processing every possible relationship between tokens, but only a small fraction of these relationships matter.

SubQ is built differently. Using a proprietary algorithm called Subquadratic Sparse Attention (SSA), it isolates the tokens and relationships that matter, ensuring compute is used efficiently.

At 2M tokens, SubQ uses 128x less compute than frontier models.

1,0080PFLOP / layer0128K512K1M2MTOKENS →31.5×64.5×128×1,008 PFLOP7.8 PFLOPFrontier Models (Dense Attention)SubQ (SSA)
Benchmarks

A leader in long-context retrieval and reasoning tasks.

Multi-Fact Retrieval

RULER · 128K multi-task retrieval99.12%

Single Fact Retrieval

Needle-in-a-haystack · 1–2M tokens100%

Single Fact Retrieval

Needle-in-a-haystack · 6–12M tokens98%
BenchmarkSubQ 1.1 SmallGPT-5.5Opus 4.8Sonnet 4.6GPT-5.4-miniGPT-5.4-nanoHaiku 4.5
Graduate-level scienceGPQA Diamond · pass@185.493.29287.587.581.767.2
Agentic financeAutomationBench13%18%16%8%0%n/r3%
Competitive programmingLiveCodeBench v6 · pass@489.79292.288.978.678.269.7

n/r = result not reported by the model provider

Technical report
Early access

Request API Access.

By submitting this form, you agree to our privacy policy and consent to receive marketing communications from Subquadratic. You can unsubscribe at any time.

Building a new class of LLMs from the architecture up.

Subquadratic is a frontier AI research company that believes the architecture layer is the highest-leverage surface in AI. While other major labs focus on incremental improvements to Transformer models, we're pushing foundational change at the model architecture level to build models and products that scale efficiently.

Built by researchers and engineers from

  • Meta
  • Google
  • Oxford
  • Cambridge
  • MIT
  • BYU
  • Amazon
  • DoorDash
  • LlamaIndex
  • TwelveLabs
  • Airbnb