More and more AI models are being used in contemporary.NET applications.
One program could be able to access:
A single application may have access to:
A high-capability model for complex reasoning
A faster model for simple requests
A lower-cost model for routine workloads
A specialized model for a particular task
A fallback model for availability problems
The engineering challenge is deciding which model should handle each request.
A static model-selection strategy is straightforward:
A routing strategy adds another decision layer:
The advantage is flexibility, but routing also introduces additional decision logic and potentially additional latency.
Microsoft.Extensions.AI provides the IChatClient
abstraction and composable chat-client pipelines, which makes it
possible to place routing, logging, retry, configuration, and other
behaviors around AI clients. The current API also provides ChatClientBuilder and DelegatingChatClient as mechanisms for composing these pipelines.
This
article explains how to benchmark a routing-based chat client against
static model selection and determine whether routing actually improves
the application's overall performance and economics.
Introduction
Suppose an application supports three request categories:
A static strategy might send every request to the same model:
A routing strategy might inspect the request and select a model:
The routing strategy can potentially improve cost, latency, or availability.
However, the router itself has a cost.
It may introduce:
Therefore, routing should be treated as an engineering hypothesis that needs to be benchmarked.
What Is Static Model Selection?
Static
model selection means the application chooses the model before
processing the request and does not dynamically change that choice.
For example:
The application knows exactly which client will process the request.
This approach is simple and predictable.
The execution path is:
What Is Routing?
Routing introduces a decision mechanism between the application and the underlying model clients.
The router can use different strategies.
For example:
Another strategy could use availability:
A third strategy could combine both:
Microsoft.Extensions.AI and IChatClient
The IChatClient
abstraction provides a common interface for chat model interactions.
This allows application code to work against an abstraction instead of
depending directly on a specific model implementation.
This is
important for benchmarking because the application can execute the same
workload against different client configurations.
For example:
A static implementation can wrap one model.
A routing implementation can select among several models.
The benchmark can then compare both using the same test scenarios.
Static Selection Architecture
A simple static architecture looks like this:
The advantage is that there is almost no selection overhead.
The main limitation is that every request follows the same model path unless application code explicitly changes the client.
Routing Architecture
A routing architecture looks like this:
The router becomes responsible for determining the target.
This can be implemented as a custom IChatClient wrapper or as a component in a chat-client pipeline.
DelegatingChatClient is specifically designed as a base type for clients that wrap another IChatClient, and the chat-client pipeline can be composed using ChatClientBuilder.Use(...).
Define the Benchmark Question
Before measuring anything, define what the benchmark is trying to prove.
For example:
Does
dynamic routing reduce cost without causing unacceptable latency or
quality degradation compared with static model selection?
That question produces several measurable dimensions:
Without
a clear hypothesis, it is easy to produce a benchmark that generates
numbers without providing an engineering conclusion.
Benchmark Scenarios
Use multiple workload categories.
Simple Requests
Examples:
Complex Requests
Examples:
Long-Context Requests
These contain larger amounts of input and can expose different model behavior.
Failure Scenarios
Simulate:
Routing should be evaluated not only when everything works but also when the preferred model fails.
Establish a Baseline
The first benchmark should use static selection.
For example:
Measure:
This becomes the baseline.
Then execute the same workload through the router.
The two measurements can then be compared.
Benchmark Harness
Create a common interface.
The request can contain:
The response can contain:
This provides a consistent measurement format.
Measure Latency
Use a monotonic timer.
This measures application-observed execution time.
Do not include unrelated operations such as loading the benchmark dataset or writing the final report inside the timed region.
Measure Routing Overhead Separately
Routing latency should not be hidden.
Consider:
If routing itself requires another model call:
That additional call can be significant.
If routing is rule-based:
The difference can be substantial.
Therefore, capture routing time independently:
Then measure the actual model request separately.
Benchmark Static Selection
A static benchmark might look like:
The exact token-usage extraction depends on the provider and client implementation.
The
benchmark should use the actual usage metadata available from the
selected client rather than estimating token counts from string length.
Benchmark Routing
A routing benchmark follows the same measurement boundary:
The important point is that both strategies receive the same benchmark request.
Rule-Based Routing
The simplest routing approach uses deterministic rules.
This has almost no classification overhead.
It is also easy to test.
The disadvantage is that rules can become increasingly complicated as workloads grow.
LLM-Based Routing
A more dynamic strategy can use a model to classify the request.
This can be flexible but introduces an additional model operation.
For example:
If static selection requires only:
the router has made the request slower.
Routing must therefore generate enough savings elsewhere to justify its own overhead.
Benchmark Routing Accuracy
A routing system should also be evaluated for decision quality.
Create an expected model category for each benchmark request:
Then compare:
Calculate:
A router that selects the wrong model frequently may not produce the expected cost or quality benefits.
Measure Model Quality
Latency alone is not sufficient.
Suppose:
Routing is faster, but the quality regression may be unacceptable.
Depending on the application, measure:
Task success rate
Structured-output validity
Groundedness
Answer relevance
Domain-specific correctness
Human evaluation
Retrieval quality for RAG workloads
The exact evaluation metric should match the application.
Cost Measurement
For each model request, capture:
Then aggregate by route.
The most useful comparison is often:
rather than simply cost per API request.
Example Cost Comparison
Imagine a benchmark produces:
| Metric | Static Selection | Routing |
|---|
| Requests | 1,000 | 1,000 |
| Success Rate | 96% | 97% |
| p50 Latency | Measure | Measure |
| p95 Latency | Measure | Measure |
| Total Tokens | Measure | Measure |
| Total Cost | Measure | Measure |
| Cost / Successful Task | Measure | Measure |
The benchmark should populate these values from actual measurements.
Avoid inserting illustrative numbers into a production recommendation unless they come from a reproducible test.
Fallback Routing
Routing can also be used for resilience.
A benchmark should measure:
For example:
The latency of the failed primary attempt should remain visible.
Otherwise, the fallback benchmark may appear faster than it actually is.
Routing and Retries
Retries introduce another variable.
Suppose the router sends a request to Model A.
The final request may succeed, but the total cost includes both failed attempts.
Track:
This provides a much more accurate picture of routing behavior.
Warm-Up Strategy
Do not use the first request as the only benchmark measurement.
Warm up each client before collecting steady-state measurements.
Then start collecting measurements.
Run cold-start tests separately if cold-start behavior matters to the production workload.
Run the Same Query Set
The static and routed systems should receive exactly the same workload.
Do not allow the router to receive easier questions than the static system.
A fixed dataset also makes regression testing easier.
Query Distribution Matters
Suppose the real application receives:
but the benchmark contains:
The resulting cost and latency numbers may not represent production.
Use a representative distribution.
If several workloads are important, benchmark them separately and report the results independently.
Concurrency Testing
A routing strategy can behave differently under load.
Test:
depending on service limits and the target workload.
Measure:
A
router that performs well at one request at a time may behave
differently when multiple requests compete for the same model capacity.
Route Distribution
Record how frequently each model is selected.
For example:
This is important for cost analysis.
If the router unexpectedly sends 80% of requests to the expensive model, the expected savings may disappear.
Static vs Routing Comparison
A useful comparison table is:
| Dimension | Static Selection | Routing |
|---|
| Implementation complexity | Low | Medium/High |
| Selection overhead | Minimal | Depends on strategy |
| Model flexibility | Low | High |
| Cost optimization | Limited | Potentially strong |
| Failover | Explicit application logic | Can be centralized |
| Debugging | Simple | More complex |
| Observability requirements | Moderate | Higher |
| Workload adaptation | Limited | Stronger |
| Predictability | High | Depends on routing policy |
Routing is not automatically better.
It is better when the additional complexity produces measurable value.
Common Benchmarking Mistakes
Comparing Different Prompts
The workload must remain consistent.
Ignoring Router Latency
A routing decision is part of the request path.
Measuring Only Average Latency
Always examine tail latency.
Ignoring Routing Accuracy
A poor route can increase cost or reduce quality.
Using Only Simple Queries
Routing benefits often appear when workloads have meaningful variation.
Ignoring Failure Paths
Fallback behavior should be benchmarked explicitly.
Ignoring Cost of Retries
A successful fallback may still have incurred multiple failed model calls.
Comparing Different Model Configurations
Keep relevant settings consistent where the benchmark is intended to isolate routing behavior.
Treating Quality as Secondary
A cheaper or faster response is not necessarily a better response.
Observability
A routing system should record enough telemetry to explain every decision.
Useful attributes include:
This allows engineers to answer:
These questions become essential when debugging production behavior.
Building a Routing Wrapper
Because DelegatingChatClient is designed for wrapping an inner IChatClient,
a custom routing abstraction can follow the same compositional pattern.
The important design choice is to keep routing policy separate from
model execution.
A simplified conceptual implementation could look like:
This is a simplified example rather than a complete implementation of a production routing client.
In a real application, the routing layer should also handle:
Cancellation
Resilience
Telemetry
Model availability
Policy validation
Error classification
Fallback
Cost tracking
Composing the Client Pipeline
The ChatClientBuilder
API supports composing intermediate chat-client stages. This allows
routing-related behavior to coexist with logging, retries, options
configuration, function invocation, and other middleware-like
components.
A conceptual pipeline can look like:
The exact ordering should be chosen deliberately.
For
example, placing telemetry around the routing layer can help measure
routing decisions separately from downstream model latency.
Release Regression Testing
Once the benchmark works, run it automatically.
For example:
The exact thresholds should be based on application requirements.
When Static Selection Is Better
Static model selection is often preferable when:
The workload is highly predictable.
One model already satisfies quality requirements.
Routing logic does not produce meaningful savings.
Simplicity is a major requirement.
The additional routing latency is unacceptable.
There are few model options.
A simpler architecture can be the better architecture.
When Routing Is Better
Routing becomes more attractive when:
Requests vary significantly in complexity.
Different models have different strengths.
Cost optimization is important.
Availability requirements justify fallback paths.
The application handles multiple workload classes.
The routing decision can be made reliably.
The operational team can observe and debug the routing behavior.
Advantages
Better Model Utilization
Different workloads can use different models.
Potential Cost Reduction
Simple requests can avoid unnecessarily expensive models.
Better Resilience
Fallback routing can improve availability when a preferred model fails.
Centralized Policy
Model-selection rules can be managed in one place.
Easier Model Evolution
New models can be introduced without rewriting every application workflow.
Disadvantages
Additional Complexity
Routing adds another component to the request path.
Routing Latency
A model-based router can add another AI operation.
Debugging Complexity
A response can depend on both the routing decision and the selected model.
More Telemetry
Engineers need visibility into route decisions, fallback, retries, and model usage.
Potential Quality Regression
An incorrect route can select a model that is cheaper or faster but less capable for the task.
Best Practices
Establish a static-model baseline before evaluating routing.
Use the same benchmark dataset for both strategies.
Measure routing latency independently.
Track p50, p95, and p99 latency.
Measure routing accuracy.
Track route distribution.
Measure token consumption and cost.
Include model quality in the benchmark.
Test fallback and retry behavior separately.
Run concurrency tests.
Use realistic production query distributions.
Keep model configuration consistent during controlled comparisons.
Record routing decisions in telemetry.
Compare cost per successful task rather than raw request cost alone.
Automate benchmark execution as part of regression testing.
Frequently Asked Questions
Is RoutingChatClient always better than static model selection?
No.
Routing introduces additional complexity and potentially additional
latency. It should be used when dynamic model selection provides
measurable value.
Does routing always reduce AI costs?
No.
A router can increase costs if it adds another model call, selects
expensive models too frequently, or causes additional retries.
Should routing be rule-based or AI-based?
Start
with deterministic rules when they are sufficient. AI-based
classification can provide more flexibility, but it introduces
additional latency and evaluation complexity.
What should I measure when benchmarking routing?
At minimum, measure latency, cost, quality, routing accuracy, success rate, fallback rate, token usage, and route distribution.
Should the router itself be included in latency?
Yes.
If the goal is to measure user-visible request latency, the routing
decision is part of the request path and should be included in total
latency. It should also be measured separately so its overhead is
visible.
How can I prove that routing is worthwhile?
Compare
routing with a static baseline using the same workload and evaluate
whether it produces an acceptable improvement in cost, latency,
reliability, or quality after accounting for routing overhead.
Conclusion
Dynamic
model routing is an attractive architecture for applications that work
with multiple AI models, but it should not be adopted simply because
multiple models are available.
The right question is whether routing produces measurable value compared with a well-defined static baseline.
A useful benchmark evaluates the complete picture:
A
routing strategy may reduce cost for simple workloads, improve
resilience during model failures, and select more capable models for
complex requests. At the same time, it can introduce classification
latency, additional operational complexity, and incorrect model
selections.
The most reliable approach is therefore empirical:
establish a static baseline, run the same workload through the routing
strategy, measure p50/p95/p99 latency, cost, quality, route accuracy,
and failure behavior, and then make the architecture decision from those
results.
In production AI systems, model routing should be
treated as a measurable optimization layer rather than an assumption
that dynamic selection is automatically better.
Best ASP.NET Core 10.0 Hosting Recommendation
One of the most important things when choosing a good ASP.NET Core 8.0 hosting is the feature and reliability.
HostForLIFE
is the leading provider of Windows hosting and affordable ASP.NET Core, their
servers are optimized for PHP web applications. The performance and the uptime of the hosting service are excellent
and the features of the web hosting plan are even greater than what many
hosting providers ask you to pay for.
At HostForLIFE.eu, customers can also experience fast ASP.NET Core
hosting. The company invested a lot of money to ensure the best and fastest
performance of the datacenters, servers, network and other facilities. Its
datacenters are equipped with the top equipments like cooling system, fire
detection, high speed Internet connection, and so on. That is why
HostForLIFEASP.NET guarantees 99.9% uptime for ASP.NET Core. And the engineers do
regular maintenance and monitoring works to assure its Orchard hosting are
security and always up.