Class JitLLMStreamingChatModel

java.lang.Object
dev.langchain4j.model.jitllm.JitLLMStreamingChatModel
All Implemented Interfaces:
StreamingChatModel, AutoCloseable

public final class JitLLMStreamingChatModel extends Object implements StreamingChatModel
A StreamingChatModel that runs a GGUF model inside the JVM with jitLLM, on a GPU through TornadoVM or on the CPU.

The model is loaded when it is built and keeps its memory (including GPU memory) until AutoCloseable.close() is called. One instance can be shared between threads, but it generates one response at a time: concurrent requests wait for each other.

StreamingChatModel.chat(ChatRequest, StreamingChatResponseHandler) blocks until the response is complete, and the handler is called on the calling thread. To avoid blocking, call it from a separate thread, for example a virtual thread.

The response is streamed token by token. When the request contains tools, the response is delivered once it is complete, because whether the model calls a tool is only known at the end. Streaming can be cancelled through the StreamingHandle passed to the handler.

Example:

try (JitLLMStreamingChatModel model = JitLLMStreamingChatModel.builder()
        .modelPath(Path.of("Qwen3-0.6B-Q8_0.gguf"))
        .build()) {
    model.chat("Tell me a story", handler);
}