Class RoutingChatModel
- All Implemented Interfaces:
ChatModel
ChatModel that sends each request to one of several chat models, as decided by a
ChatModelRouter. It can be used wherever a ChatModel is expected (AI Services, agents, RAG), for
example to send simple requests to a small, cheap model and complex ones to a larger model:
ChatModel chatModel = RoutingChatModel.builder()
.route("simple", "Greetings, short factual questions, simple lookups", smallModel)
.route("complex", "Multi-step reasoning, code, analysis", largeModel)
.router(new DecisionModelChatModelRouter(decisionModel))
.defaultRoute("complex")
.build();
The selected model handles the request as if it was called directly: its default parameters and listeners apply.
Since the request can go to models of different providers, set only common parameters on requests
(ChatRequestParameters), not provider-specific ones.
The name of the selected route is:
- added to the listener attributes of the call, under
ROUTE_ATTRIBUTE, so that the listeners of the selected model can report it; - stored in the
attributesof the returnedAiMessage, underROUTE_ATTRIBUTE.
AiMessage that requested the tools, without asking the router. This also works when the
messages are stored in a persistent chat memory, as long as it keeps the attributes of the messages.
supportedCapabilities() returns the capabilities supported by at least one route. A request that needs a
capability (such as a JSON schema response format) is only routed to the routes that declare it: the router only
sees those routes, and the default route is replaced by the first of them if it does not declare the capability.
If no route declares the capability, all routes remain candidates, and the selected model accepts or rejects the
request itself, as when it is called directly.
The routing chat model has no default request parameters of its own:
ChatModel.defaultRequestParameters() returns empty parameters, and the default parameters of the selected model
apply to each request.
The asynchronous method (ChatModel.chatAsync(ChatRequest)) selects the route with
ChatModelRouter.routeAsync(ChatModelRoutingRequest), so it never blocks. A router that does not implement it,
such as a router written as a lambda, fails the call with an
AsyncNotSupportedException.
- Since:
- 1.21.0
- See Also:
-
Nested Class Summary
Nested Classes -
Field Summary
Fields -
Constructor Summary
Constructors -
Method Summary
Modifier and TypeMethodDescriptionstatic RoutingChatModel.Builderbuilder()chat(ChatRequest chatRequest, ChatRequestOptions options) Sends a chat request with additional invocation options.chatAsync(ChatRequest chatRequest, ChatRequestOptions options) Sends a non-blocking chat request with additional invocation options.The name of the default route, used when the router does not select one.doChat(ChatRequest chatRequest) doChatAsync(ChatRequest chatRequest) SPI hook for a genuinely non-blocking chat implementation, invoked byChatModel.chatAsync(ChatRequest).routes()The routes, in the order in which they were configured.
-
Field Details
-
ROUTE_ATTRIBUTE
The key under which the name of the selected route is stored in the listener attributes of the call and in the attributes of the returnedAiMessage. Since the attribute is stored in chat memories together with the message, the key and the value (the name of the route) are part of the persisted data: keep route names stable. When routing chat models are nested, the outer one overwrites the value of the inner one.- See Also:
-
-
Constructor Details
-
RoutingChatModel
-
-
Method Details
-
chat
Description copied from interface:ChatModelSends a chat request with additional invocation options.- Specified by:
chatin interfaceChatModel- Parameters:
chatRequest- aChatRequest, containing all the inputs to the LLMoptions- aChatRequestOptionscarrying listener attributes and other per-call metadata- Returns:
- a
ChatResponse, containing all the outputs from the LLM
-
doChat
-
chatAsync
public CompletableFuture<ChatResponse> chatAsync(ChatRequest chatRequest, ChatRequestOptions options) Description copied from interface:ChatModelSends a non-blocking chat request with additional invocation options.- Specified by:
chatAsyncin interfaceChatModel- Parameters:
chatRequest- aChatRequest, containing all the inputs to the LLMoptions- aChatRequestOptionscarrying listener attributes and other per-call metadata- Returns:
- a
CompletableFutureof theChatResponse - See Also:
-
doChatAsync
Description copied from interface:ChatModelSPI hook for a genuinely non-blocking chat implementation, invoked byChatModel.chatAsync(ChatRequest).The default returns a failed future carrying
AsyncNotSupportedExceptionto signal that this model has no native asynchronous implementation. Callers on the asynchronous and reactive path (for example the non-blocking RAG stages) detect this and either offload the blockingChatModel.doChat(ChatRequest)or fail loudly with an actionable message. A model backed by remote HTTP I/O overrides this with a genuinely asynchronous call (no thread parked).- Specified by:
doChatAsyncin interfaceChatModel- Parameters:
chatRequest- aChatRequest, containing all the inputs to the LLM- Returns:
- a
CompletableFutureof theChatResponse
-
supportedCapabilities
- Specified by:
supportedCapabilitiesin interfaceChatModel
-
routes
The routes, in the order in which they were configured. -
defaultRoute
The name of the default route, used when the router does not select one. -
builder
-