Multi-Model Orchestration: When You Need More Than One LLM per Task
The default instinct when a single-model approach has a problem is to add a second model. The document extraction is not accurate enough, so add a verification model to check the first one. The classification model is making errors, so add a judge model to review borderline cases. This instinct is sometimes right and often wrong. Adding a second model adds latency, adds cost, adds a new failure mode, and only helps if the second model is actually better at the specific task you are using it for than the first. Sometimes the second model clearly justifies itself. Sometimes a better-prompted single model does the job. Here is how to tell the difference before you build.
When multiple models are clearly the right call
Three scenarios where multi-model orchestration justifies itself:
Different models for different modalities. A document processing pipeline that needs to handle both scanned PDFs (requiring OCR + image understanding) and structured data exports (requiring text analysis) benefits from using a vision-capable model for the former and a text-only model for the latter. Claude 3.5 Haiku for structured text processing at $0.0008/k input tokens, Claude 3.5 Sonnet for PDFs requiring image understanding at $0.003/k. Using Sonnet for everything is three times more expensive for the text-only cases with no accuracy benefit.
Cost-tiered routing for confidence thresholds. A classification system where 80% of cases are straightforward can use a fast, cheap model (Haiku, GPT-4o mini) for the easy cases and route only low-confidence cases to a more capable model. The arithmetic is simple. If 80% of records go to a model that costs a tenth as much and 20% go to the premium model, the blended cost is 28% of sending everything to the premium model. That is a 72% saving, provided accuracy on the easy cases holds, which you should measure.
Specialized models for specific subtasks. When a subtask has a specialized model that really does outperform general models on it (a fine-tuned model for a highly domain-specific classification problem, or an embedding model optimized for a specific retrieval task), using the specialized model for that subtask and a general model for the rest is legitimate. The key word is "genuinely outperforms." Benchmark the specialized model against the general model on your actual data before committing to the added complexity.
When the second model does not help
The common failure pattern: the first model makes errors. Someone proposes adding a second model to verify the first. The second model is the same capability tier as the first, often the same model. It makes different errors, not fewer errors. The combined system has the errors of both models, higher latency, and higher cost.
A verification model only helps when it is meaningfully better at detecting the specific errors the primary model makes. A Claude 3.5 Sonnet verification pass on Claude 3.5 Sonnet extraction output typically does not catch the extraction errors. If the primary model was wrong, the verification model usually agrees with it because they share similar reasoning patterns and will make similar mistakes on ambiguous cases.
Before adding a verification model, test whether it catches the errors you want caught. Take a sample of cases where the primary model was wrong and run the verification model on them. If it gets most of those wrong too, it is adding latency and cost for little benefit. A better-prompted single model often beats a two-model verification setup.
The orchestration complexity cost
Multi-model systems have a failure surface proportional to the number of models they use. Each model API call can fail, time out, or return unexpected output. Each model transition is a parsing step where the output of one model needs to be correctly interpreted by the next. Each additional model version update (a provider updates Claude Sonnet, GPT-4o, Gemini, and they all behave slightly differently afterward) can change system behavior in ways that require re-testing and re-calibration.
The incidents look like this: a rate limit that applies to the verification model but not the primary one, so the primary runs, the verification does not, and outputs go out with no quality check; a model update that changes the output format of a middle-step model and breaks the parser feeding the next step; and a cold-start latency spike on a specialized model that pushes the total request time past the SLA.
Our take
The routing architecture that works
The pattern with the best cost-accuracy tradeoff: a fast, cheap routing classifier followed by capability-matched processing.
The routing classifier is a simple LLM call (or a classifier model) that assigns each input to a tier: routine (high confidence, standard format), moderate (needs care, some ambiguity), or complex (low confidence, unusual format, high-stakes). Routine goes to the fast model. Moderate goes to the standard model. Complex goes to the most capable model and may include a human-in-the-loop step.
The routing classifier needs to be accurate, but the cost of a misrouting is usually just cost or latency, not correctness, because each tier is still capable of handling any input, just at different cost and quality levels. Build and test the routing classifier as carefully as the processing models. A bad routing classifier that sends everything to the expensive tier is worse than no routing at all.