One family. Three models. Built for every scale.
Choose the model that matches the job — not a one-size-fits-all bill. Magma for high-stakes reasoning, Blaze for everyday production agents, Ember for high-volume long-context work. Same platform. Clear price points. Built for IDE, Cloud, CLI, and API.
Flagship reasoning for work where getting it wrong is expensive.
Magma is 4RGED’s highest-capability model. It is built for long-horizon plans, hard trade-offs, and multi-step agent runs where depth matters more than raw speed.
Use Magma when the model must hold a large design space in mind — system boundaries, migration risks, policy constraints — and still produce a plan you can execute.
It excels at architecture reviews, complex refactors, and agent loops that call tools across many files or services before committing to an answer.
If a cheaper model keeps missing edge cases or collapsing plans too early, Magma is the upgrade path inside the same 4RGED family.
Input context
500K
Output context
500K
Latency
Standard
Price in/out
$4.00 / $8.00
What it does well
Frontier-depth reasoning
Handles ambiguous requirements, conflicting constraints, and multi-hop decisions without shortcutting to a shallow answer.
Architecture & planning
Breaks large changes into sequenced steps, surfaces risks early, and keeps system-wide context across long sessions.
Multi-step agent execution
Stays coherent across tool calls, retries, and verification loops — built for agents that must finish the job, not just draft a reply.
High-stakes review quality
Stronger at catching policy, security, and correctness issues before they land in production.
Choose this when
Complex reasoning, architecture decisions, and critical agent workloads.
- The cost of a wrong architectural choice outweighs Magma’s higher unit price.
- You need deep, multi-file or multi-service plans — not a single-file patch.
- Your agents run long tool chains and must preserve intent across many steps.
- You are reviewing security-sensitive or regulated changes.
Analytics
Magma analytics
Indexed scores are relative within the 4RGED family and peer set. Prices are USD per 1M tokens.
Context
500K
In / out
Blended list
$6.00
Avg of in+out /1M
Output vs Opus
3.1×
Lower output cost
Latency class
Std
Depth over speed
Peer price / 1M tokens
Profile radar
Capability profile
Index 0–100- Reasoning98
- Coding92
- Agents96
- Speed64
- Context72
- Cost efficiency86
Customer use cases
System design & architecture reviews
Evaluate trade-offs across services, data stores, and rollout plans with explicit risks and alternatives.
Long-horizon coding agents
Drive agents that research, edit, test, and revise across a large codebase before proposing a merge-ready change.
Hard refactors & migrations
Plan and execute migrations where order-of-operations and compatibility matter as much as the code itself.
Policy, security & correctness review
Pressure-test PRs and agent output for subtle failures, unsafe patterns, and compliance gaps.
API · 4rge-magma
The production default — strong coding agents at a lower unit cost.
Blaze is the workhorse of the 4RGED family: excellent coding and agent quality with enough headroom for real production systems — without Magma pricing on every request.
Most teams should start here. Blaze covers pair-programming in the IDE, Cloud parallel runs, CI fixes, and CLI automation with reliable reasoning and fast turnaround.
It balances depth and cost: strong enough for production agent orchestration, efficient enough to run continuously across a team or fleet.
Move up to Magma only when Blaze consistently under-delivers on hard planning; move down to Ember when the job is high-volume summarization or routing.
Input context
1M
Output context
1M
Latency
Fast
Price in/out
$2.00 / $4.00
What it does well
Production coding quality
Writes, edits, and explains code with strong repo awareness — tuned for IDE pair-programming and review loops.
Reliable under load
Built for everyday production traffic: consistent behavior across many sessions, orgs, and agent types.
Agentic orchestration
Coordinates tools, tests, and multi-agent workflows without needing flagship pricing on every turn.
Fast enough for interactive work
Responsive for IDE and Cloud sessions where engineers are waiting on the next edit or answer.
Choose this when
Production workloads, coding agents, and day-to-day orchestration.
- You want one default model for IDE, Cloud, and CI agents.
- Coding quality matters, but Magma-level depth is not required on every call.
- You run many parallel agent sessions and need predictable cost.
- You need a million-token context window for real repositories and tickets.
Analytics
Blaze analytics
Indexed scores are relative within the 4RGED family and peer set. Prices are USD per 1M tokens.
Context
1M
In / out
Blended list
$3.00
Avg of in+out /1M
Output vs Sonnet
3.8×
Lower output cost
Latency class
Fast
Interactive default
Peer price / 1M tokens
Profile radar
Capability profile
Index 0–100- Reasoning86
- Coding94
- Agents90
- Speed88
- Context88
- Cost efficiency92
Customer use cases
IDE pair-programming agents
Inline edits, explanations, and multi-file changes while an engineer stays in flow.
Cloud parallel agent fleets
Spin up many agents on tickets, PRs, or migrations without flagship cost on each run.
CI, CLI & automation
Fix failing checks, generate patches, and drive scripted workflows from the terminal or pipelines.
Team default for new projects
Standardize on one model so policy, prompts, and evals stay simple across the org.
API · 4rge-blaze
Maximum context and throughput for everyday high-volume work.
Ember is optimized for scale: the largest context window in the family at the lowest price — ideal when you need to read a lot, respond quickly, and run many jobs.
Reach for Ember when the bottleneck is volume or document size — chat assistants, summarization, classification, routing, and ingestion pipelines — not deep architecture planning.
Its 2M context window lets you keep long threads, large PDFs, or multi-file dumps in a single pass instead of brittle chunking.
Pair Ember with Blaze or Magma in a router: Ember handles the high-volume front door; escalate harder tasks to Blaze or Magma only when needed.
Input context
2M
Output context
2M
Latency
Fastest
Price in/out
$1.00 / $2.00
What it does well
High-throughput responses
Built for queues and interactive chat where latency and cost per request dominate.
Large document & dataset intake
Ingest long docs, transcripts, and dumps in fewer passes — less glue code, fewer missed sections.
Efficient long-context processing
2M input and output context for jobs that would otherwise require aggressive truncation.
Smart routing front door
Classify and escalate: Ember handles volume; Magma or Blaze take the hard cases.
Choose this when
Everyday chat, summaries, triage, and high-volume long-context jobs.
- You process large documents or long conversations in a single request.
- Unit economics matter more than maximum reasoning depth.
- You need the fastest turnaround in the 4RGED family.
- You want a cheap first-pass model before escalating to Blaze or Magma.
Analytics
Ember analytics
Indexed scores are relative within the 4RGED family and peer set. Prices are USD per 1M tokens.
Context
2M
Largest in family
Blended list
$1.50
Avg of in+out /1M
Output vs Haiku
2.5×
Lower output cost
Latency class
Fastest
High-volume ready
Peer price / 1M tokens
Profile radar
Capability profile
Index 0–100- Reasoning72
- Coding70
- Agents68
- Speed98
- Context98
- Cost efficiency96
Customer use cases
Customer & internal chat assistants
Fast answers across long histories without burning Magma or Blaze budget on every message.
Summaries & briefings
Collapse tickets, meetings, and docs into actionable notes at high volume.
Classification & triage pipelines
Label, route, and prioritize incoming work before a deeper model gets involved.
Large-context ingestion
Pull big files and datasets into agents for search, extract, and transform workflows.
API · 4rge-ember
List prices for input and output. Magma, Blaze, and Ember stay below typical frontier output rates while keeping production-ready context windows. Select a 4RGED row to jump back to that model’s full story.
Input tokens (per 1M)
Output tokens (per 1M)
Select a 4RGED row to jump focus back to that model above.
0x
- Built for frontier-scale workloads.
- Massive context windows.
- Fewer inference passes.
- Lower real-world compute cost.
Save up to 4x on input & output tokens.
Faster responses with optimized inference.
Secure. Scalable. Reliable.
Built for builders. APIs, agents & real-world scale.
