Class JitLLMStreamingChatModel
- All Implemented Interfaces:
StreamingChatModel, AutoCloseable
StreamingChatModel that runs a GGUF model inside the JVM with jitLLM,
on a GPU through TornadoVM or on the CPU.
The model is loaded when it is built and keeps its memory (including GPU memory) until AutoCloseable.close() is called.
One instance can be shared between threads, but it generates one response at a time:
concurrent requests wait for each other.
StreamingChatModel.chat(ChatRequest, StreamingChatResponseHandler) blocks until the response is complete,
and the handler is called on the calling thread. To avoid blocking, call it from a separate thread,
for example a virtual thread.
The response is streamed token by token. When the request contains tools, the response is
delivered once it is complete, because whether the model calls a tool is only known at the end.
Streaming can be cancelled through the StreamingHandle
passed to the handler.
Example:
try (JitLLMStreamingChatModel model = JitLLMStreamingChatModel.builder()
.modelPath(Path.of("Qwen3-0.6B-Q8_0.gguf"))
.build()) {
model.chat("Tell me a story", handler);
}
-
Nested Class Summary
Nested Classes -
Method Summary
Modifier and TypeMethodDescriptionbuilder()Creates a new builder.voidclose()Releases the memory held by the model, including the device memory when running on a GPU.Returns the default request parameters of this model.voiddoChat(ChatRequest chatRequest, StreamingChatResponseHandler handler) Returns the listeners of this model.Methods inherited from class Object
clone, equals, finalize, getClass, hashCode, notify, notifyAll, toString, wait, wait, waitMethods inherited from interface StreamingChatModel
chat, chat, chat, chat, chat, chat, chat, chat, defaultRequestParameters, doChat, listeners, provider, supportedCapabilities
-
Method Details
-
builder
Creates a new builder.- Returns:
- a new builder
-
doChat
- Specified by:
doChatin interfaceStreamingChatModel
-
defaultRequestParameters
Returns the default request parameters of this model.- Returns:
- the default request parameters
-
listeners
Returns the listeners of this model.- Returns:
- the listeners
-
close
public void close()Releases the memory held by the model, including the device memory when running on a GPU. Waits for a running generation to finish first. After the model is closed, every request fails with anIllegalStateException.- Specified by:
closein interfaceAutoCloseable
-