Class RoutingStreamingChatModel
- All Implemented Interfaces:
StreamingChatModel
RoutingChatModel: a StreamingChatModel that sends each request to one
of several streaming chat models, as decided by a ChatModelRouter.
StreamingChatModel streamingChatModel = RoutingStreamingChatModel.builder()
.route("simple", "Greetings, short factual questions, simple lookups", smallStreamingModel)
.route("complex", "Multi-step reasoning, code, analysis", largeStreamingModel)
.router(new DecisionModelChatModelRouter(decisionModel))
.defaultRoute("complex")
.build();
See RoutingChatModel for how requests are routed and where the selected route is recorded (with
chat(ChatRequest), which takes no options, only in the attributes of the response).
With a StreamingChatResponseHandler, the route is selected on the calling thread with
ChatModelRouter.route(ChatModelRoutingRequest), before streaming starts. This blocks the calling thread
until the route is selected (with DecisionModelChatModelRouter, until the decision model answers), so do
not call it from an event loop thread. The non-blocking
chat(ChatRequest) returning a Flow.Publisher selects it with
ChatModelRouter.routeAsync(ChatModelRoutingRequest) instead: a router that does not implement it, such as a
router written as a lambda, fails the stream with an AsyncNotSupportedException.
- Since:
- 1.21.0
-
Nested Class Summary
Nested Classes -
Constructor Summary
ConstructorsModifierConstructorDescriptionprotected -
Method Summary
Modifier and TypeMethodDescriptionbuilder()chat(ChatRequest request) Reactive entry point: sends a chat request and returns aFlow.PublisherofChatModelStreamingEvents.voidchat(ChatRequest request, ChatRequestOptions options, StreamingChatResponseHandler handler) Sends a streaming chat request with additional invocation options.The name of the default route, used when the router does not select one.doChat(ChatRequest chatRequest) Provider-specific implementation of the reactive stream returned byStreamingChatModel.chat(ChatRequest)(which wraps it withChatModelListenerinvocation).voiddoChat(ChatRequest chatRequest, StreamingChatResponseHandler handler) routes()The routes, in the order in which they were configured.Methods inherited from class Object
clone, equals, finalize, getClass, hashCode, notify, notifyAll, toString, wait, wait, waitMethods inherited from interface StreamingChatModel
chat, chat, chat, chat, chat, chat, defaultRequestParameters, listeners, provider
-
Constructor Details
-
RoutingStreamingChatModel
-
-
Method Details
-
chat
public void chat(ChatRequest request, ChatRequestOptions options, StreamingChatResponseHandler handler) Description copied from interface:StreamingChatModelSends a streaming chat request with additional invocation options.- Specified by:
chatin interfaceStreamingChatModel- Parameters:
request- aChatRequest, containing all the inputs to the LLMoptions- aChatRequestOptionscarrying listener attributes and other per-call metadatahandler- aStreamingChatResponseHandlerthat will handle streaming response from the LLM
-
doChat
- Specified by:
doChatin interfaceStreamingChatModel
-
chat
Description copied from interface:StreamingChatModelReactive entry point: sends a chat request and returns aFlow.PublisherofChatModelStreamingEvents.The publisher is cold: nothing happens until you subscribe, and each
subscribe()call initiates a new LLM request. It emits, in this order (eachCompleteToolCallarriving as soon as that tool call finishes assembling, so it interleaves with the next call's chunks):- 0..N
PartialThinking(thinking/reasoning chunks), - 0..N
PartialResponse(text chunks), - 0..N
PartialToolCall(tool-call argument chunks), - 0..N
CompleteToolCall(assembled tool calls), - 0..N
RawStreamingEvent(provider-specific raw events, interleaved with the above), - exactly one terminal
CompleteResponse(wrapping the aggregated finalChatResponse),
onComplete. On failure,onErroris signaled afteronSubscribe.Registered
ChatModelListeners are invoked:onRequeston each new subscription (just before the underlying request goes out),onResponseafter the terminalChatResponseis emitted,onErroron failure.If the
Flow.Subscriberthrows fromonNext(or any other signal method), it violates the Reactive Streams contract (Rule 2.13): the stream is cancelled and no further events are delivered, and noChatModelListenercallback fires for it — neitheronResponsenoronError. This differs from the handler-basedStreamingChatModel.chat(ChatRequest, StreamingChatResponseHandler)path, which catches exceptions thrown from handler callbacks, reports them toonError, and keeps streaming.Subscribers must be prepared to receive
ChatModelStreamingEventsubtypes they do not recognize and ignore them. New event types may be introduced over time (and providers may surface unmapped events asRawStreamingEvent), so consuming this stream with an exhaustive type switch that lacks a default branch is unsafe.Demand and back-pressure. This streams a finite, bounded-rate source — an LLM response over HTTP. Implementations are not required to propagate subscriber demand to the model: meaningfully throttling an LLM is impractical (its work and cost are incurred regardless of how fast the response is read, and stalling the transport to slow it down only risks provider/proxy idle timeouts). An implementation therefore typically consumes the response eagerly and relays events through a bounded internal buffer. A subscriber that requests fewer items than are produced may thus cause buffering and, once the buffer is exhausted, a terminal error. Subscribers should request liberally (e.g.
Long.MAX_VALUE) and must not block or perform heavy work inonNext— offload it to another thread.Threading. Events are delivered on the model's own thread — for HTTP models, the transport's I/O worker that reads the response (the JDK HTTP client's
HttpClient-*workers), the same scarce, shared threads aChatModelListenercallback runs on. Blocking there stalls this stream and, under concurrency, degrades throughput for every in-flight call.- Specified by:
chatin interfaceStreamingChatModel
- 0..N
-
doChat
Description copied from interface:StreamingChatModelProvider-specific implementation of the reactive stream returned byStreamingChatModel.chat(ChatRequest)(which wraps it withChatModelListenerinvocation). Implementations must honor the event ordering and the demand / back-pressure expectations documented onStreamingChatModel.chat(ChatRequest)— in particular, they typically consume the response eagerly and relayChatModelStreamingEvents through a bounded buffer rather than propagating subscriber demand to the model.The default implementation returns an immediately-failing Publisher carrying
AsyncNotSupportedExceptionto signal that this model has no native reactive-streaming implementation; a provider that does not support reactive streaming leaves it unimplemented (consistent withChatModel#doChatAsyncand the other async defaults).- Specified by:
doChatin interfaceStreamingChatModel
-
supportedCapabilities
- Specified by:
supportedCapabilitiesin interfaceStreamingChatModel
-
routes
The routes, in the order in which they were configured. -
defaultRoute
The name of the default route, used when the router does not select one. -
builder
-