Class RoutingStreamingChatModel

java.lang.Object
dev.langchain4j.model.chat.router.RoutingStreamingChatModel
All Implemented Interfaces:
StreamingChatModel

@Experimental public class RoutingStreamingChatModel extends Object implements StreamingChatModel
The streaming counterpart of RoutingChatModel: a StreamingChatModel that sends each request to one of several streaming chat models, as decided by a ChatModelRouter.
StreamingChatModel streamingChatModel = RoutingStreamingChatModel.builder()
        .route("simple", "Greetings, short factual questions, simple lookups", smallStreamingModel)
        .route("complex", "Multi-step reasoning, code, analysis", largeStreamingModel)
        .router(new DecisionModelChatModelRouter(decisionModel))
        .defaultRoute("complex")
        .build();
See RoutingChatModel for how requests are routed and where the selected route is recorded (with chat(ChatRequest), which takes no options, only in the attributes of the response).

With a StreamingChatResponseHandler, the route is selected on the calling thread with ChatModelRouter.route(ChatModelRoutingRequest), before streaming starts. This blocks the calling thread until the route is selected (with DecisionModelChatModelRouter, until the decision model answers), so do not call it from an event loop thread. The non-blocking chat(ChatRequest) returning a Flow.Publisher selects it with ChatModelRouter.routeAsync(ChatModelRoutingRequest) instead: a router that does not implement it, such as a router written as a lambda, fails the stream with an AsyncNotSupportedException.

Since:
1.21.0
  • Constructor Details

  • Method Details

    • chat

      public void chat(ChatRequest request, ChatRequestOptions options, StreamingChatResponseHandler handler)
      Description copied from interface: StreamingChatModel
      Sends a streaming chat request with additional invocation options.
      Specified by:
      chat in interface StreamingChatModel
      Parameters:
      request - a ChatRequest, containing all the inputs to the LLM
      options - a ChatRequestOptions carrying listener attributes and other per-call metadata
      handler - a StreamingChatResponseHandler that will handle streaming response from the LLM
    • doChat

      public void doChat(ChatRequest chatRequest, StreamingChatResponseHandler handler)
      Specified by:
      doChat in interface StreamingChatModel
    • chat

      Description copied from interface: StreamingChatModel
      Reactive entry point: sends a chat request and returns a Flow.Publisher of ChatModelStreamingEvents.

      The publisher is cold: nothing happens until you subscribe, and each subscribe() call initiates a new LLM request. It emits, in this order (each CompleteToolCall arriving as soon as that tool call finishes assembling, so it interleaves with the next call's chunks):

      followed by onComplete. On failure, onError is signaled after onSubscribe.

      Registered ChatModelListeners are invoked: onRequest on each new subscription (just before the underlying request goes out), onResponse after the terminal ChatResponse is emitted, onError on failure.

      If the Flow.Subscriber throws from onNext (or any other signal method), it violates the Reactive Streams contract (Rule 2.13): the stream is cancelled and no further events are delivered, and no ChatModelListener callback fires for it — neither onResponse nor onError. This differs from the handler-based StreamingChatModel.chat(ChatRequest, StreamingChatResponseHandler) path, which catches exceptions thrown from handler callbacks, reports them to onError, and keeps streaming.

      Subscribers must be prepared to receive ChatModelStreamingEvent subtypes they do not recognize and ignore them. New event types may be introduced over time (and providers may surface unmapped events as RawStreamingEvent), so consuming this stream with an exhaustive type switch that lacks a default branch is unsafe.

      Demand and back-pressure. This streams a finite, bounded-rate source — an LLM response over HTTP. Implementations are not required to propagate subscriber demand to the model: meaningfully throttling an LLM is impractical (its work and cost are incurred regardless of how fast the response is read, and stalling the transport to slow it down only risks provider/proxy idle timeouts). An implementation therefore typically consumes the response eagerly and relays events through a bounded internal buffer. A subscriber that requests fewer items than are produced may thus cause buffering and, once the buffer is exhausted, a terminal error. Subscribers should request liberally (e.g. Long.MAX_VALUE) and must not block or perform heavy work in onNext — offload it to another thread.

      Threading. Events are delivered on the model's own thread — for HTTP models, the transport's I/O worker that reads the response (the JDK HTTP client's HttpClient-* workers), the same scarce, shared threads a ChatModelListener callback runs on. Blocking there stalls this stream and, under concurrency, degrades throughput for every in-flight call.

      Specified by:
      chat in interface StreamingChatModel
    • doChat

      public Flow.Publisher<ChatModelStreamingEvent> doChat(ChatRequest chatRequest)
      Description copied from interface: StreamingChatModel
      Provider-specific implementation of the reactive stream returned by StreamingChatModel.chat(ChatRequest) (which wraps it with ChatModelListener invocation). Implementations must honor the event ordering and the demand / back-pressure expectations documented on StreamingChatModel.chat(ChatRequest) — in particular, they typically consume the response eagerly and relay ChatModelStreamingEvents through a bounded buffer rather than propagating subscriber demand to the model.

      The default implementation returns an immediately-failing Publisher carrying AsyncNotSupportedException to signal that this model has no native reactive-streaming implementation; a provider that does not support reactive streaming leaves it unimplemented (consistent with ChatModel#doChatAsync and the other async defaults).

      Specified by:
      doChat in interface StreamingChatModel
    • supportedCapabilities

      public Set<Capability> supportedCapabilities()
      Specified by:
      supportedCapabilities in interface StreamingChatModel
    • routes

      public List<ChatModelRoute> routes()
      The routes, in the order in which they were configured.
    • defaultRoute

      public String defaultRoute()
      The name of the default route, used when the router does not select one.
    • builder

      public static RoutingStreamingChatModel.Builder builder()