Class JitLLMChatModel

java.lang.Object
dev.langchain4j.model.jitllm.JitLLMChatModel
All Implemented Interfaces:
ChatModel, AutoCloseable

public final class JitLLMChatModel extends Object implements ChatModel
A ChatModel that runs a GGUF model inside the JVM with jitLLM, on a GPU through TornadoVM or on the CPU.

The model is loaded when it is built and keeps its memory (including GPU memory) until AutoCloseable.close() is called. One instance can be shared between threads, but it generates one response at a time: concurrent requests wait for each other.

Example:

try (JitLLMChatModel model = JitLLMChatModel.builder()
        .modelPath(Path.of("Qwen3-0.6B-Q8_0.gguf"))
        .build()) {
    String answer = model.chat("What is the capital of Germany?");
}
  • Method Details

    • builder

      public static JitLLMChatModel.JitLLMChatModelBuilder builder()
      Creates a new builder.
      Returns:
      a new builder
    • doChat

      public ChatResponse doChat(ChatRequest chatRequest)
      Specified by:
      doChat in interface ChatModel
    • defaultRequestParameters

      public ChatRequestParameters defaultRequestParameters()
      Returns the default request parameters of this model.
      Returns:
      the default request parameters
    • listeners

      public List<ChatModelListener> listeners()
      Returns the listeners of this model.
      Returns:
      the listeners
    • close

      public void close()
      Releases the memory held by the model, including the device memory when running on a GPU. Waits for a running generation to finish first. After the model is closed, every request fails with an IllegalStateException.
      Specified by:
      close in interface AutoCloseable