Class JitLLMChatModel
java.lang.Object
dev.langchain4j.model.jitllm.JitLLMChatModel
- All Implemented Interfaces:
ChatModel, AutoCloseable
A
ChatModel that runs a GGUF model inside the JVM with jitLLM,
on a GPU through TornadoVM or on the CPU.
The model is loaded when it is built and keeps its memory (including GPU memory) until AutoCloseable.close() is called.
One instance can be shared between threads, but it generates one response at a time:
concurrent requests wait for each other.
Example:
try (JitLLMChatModel model = JitLLMChatModel.builder()
.modelPath(Path.of("Qwen3-0.6B-Q8_0.gguf"))
.build()) {
String answer = model.chat("What is the capital of Germany?");
}
-
Nested Class Summary
Nested Classes -
Method Summary
Modifier and TypeMethodDescriptionbuilder()Creates a new builder.voidclose()Releases the memory held by the model, including the device memory when running on a GPU.Returns the default request parameters of this model.doChat(ChatRequest chatRequest) Returns the listeners of this model.
-
Method Details
-
builder
Creates a new builder.- Returns:
- a new builder
-
doChat
-
defaultRequestParameters
Returns the default request parameters of this model.- Returns:
- the default request parameters
-
listeners
Returns the listeners of this model.- Returns:
- the listeners
-
close
public void close()Releases the memory held by the model, including the device memory when running on a GPU. Waits for a running generation to finish first. After the model is closed, every request fails with anIllegalStateException.- Specified by:
closein interfaceAutoCloseable
-