Class JitLLMStreamingChatModel.JitLLMStreamingChatModelBuilder
java.lang.Object
dev.langchain4j.model.jitllm.JitLLMStreamingChatModel.JitLLMStreamingChatModelBuilder
- Enclosing class:
JitLLMStreamingChatModel
Builder for
JitLLMStreamingChatModel.-
Method Summary
Modifier and TypeMethodDescriptionbuild()Builds the model.contextLength(Integer contextLength) Sets the context window: the maximum number of tokens of the whole conversation (system message, history, tool definitions and the generated response).defaultRequestParameters(ChatRequestParameters defaultRequestParameters) Sets the parameters used for every request, unless the request sets them itself.listeners(List<ChatModelListener> listeners) Sets the listeners notified about every request, response and error.Sets the maximum number of tokens to generate per response, including the thinking.Sets the path to the model file in GGUF format.Sets whether the model runs on a GPU (true) or on the CPU (false).returnThinking(Boolean returnThinking) Sets whether the thinking of reasoning models (the text between<think>and</think>) is returned inAiMessage.thinking()and streamed toStreamingChatResponseHandler.onPartialThinking(PartialThinking).Sets the seed of the random sampling, to make responses reproducible.stopSequences(List<String> stopSequences) Sets the sequences that end the response when they appear in the answer.temperature(Double temperature) Sets the sampling temperature.Sets whether reasoning models think before answering:trueenables thinking,falsedisables it, and when not set, the model's own default applies (Qwen 3, for example, thinks).Sets the nucleus sampling probability.
-
Method Details
-
modelPath
Sets the path to the model file in GGUF format. Required.- Parameters:
modelPath- the path to the GGUF file- Returns:
- this builder
-
contextLength
public JitLLMStreamingChatModel.JitLLMStreamingChatModelBuilder contextLength(Integer contextLength) Sets the context window: the maximum number of tokens of the whole conversation (system message, history, tool definitions and the generated response). Memory for the whole window is reserved when the model is loaded. Default: 4096.- Parameters:
contextLength- the context window in tokens- Returns:
- this builder
-
onGPU
Sets whether the model runs on a GPU (true) or on the CPU (false).Running on a GPU requires the JVM to be started through TornadoVM with
-Duse.tornadovm=true, otherwise building the model fails with anIllegalStateException. The GPU backend (CUDA, OpenCL or Metal) is the one provided by the installed TornadoVM SDK.Default:
trueif the JVM was started with-Duse.tornadovm=true,falseotherwise.- Parameters:
onGPU- whether to run on a GPU- Returns:
- this builder
-
temperature
Sets the sampling temperature. Default: 0.1.- Parameters:
temperature- the temperature- Returns:
- this builder
-
topP
Sets the nucleus sampling probability. Default: 0.95.- Parameters:
topP- the nucleus sampling probability- Returns:
- this builder
-
maxTokens
Sets the maximum number of tokens to generate per response, including the thinking. Default: 512.- Parameters:
maxTokens- the maximum number of tokens to generate- Returns:
- this builder
-
stopSequences
public JitLLMStreamingChatModel.JitLLMStreamingChatModelBuilder stopSequences(List<String> stopSequences) Sets the sequences that end the response when they appear in the answer. The response is cut before the sequence. The thinking of reasoning models is not checked.- Parameters:
stopSequences- the stop sequences- Returns:
- this builder
-
seed
Sets the seed of the random sampling, to make responses reproducible. Default: a different seed for every request.- Parameters:
seed- the seed- Returns:
- this builder
-
think
Sets whether reasoning models think before answering:trueenables thinking,falsedisables it, and when not set, the model's own default applies (Qwen 3, for example, thinks). Models that cannot think ignore this setting. Thinking improves answers to complex questions, but the thinking tokens count againstmaxTokens(Integer)and take time to generate.- Parameters:
think- whether the model thinks before answering- Returns:
- this builder
-
returnThinking
public JitLLMStreamingChatModel.JitLLMStreamingChatModelBuilder returnThinking(Boolean returnThinking) Sets whether the thinking of reasoning models (the text between<think>and</think>) is returned inAiMessage.thinking()and streamed toStreamingChatResponseHandler.onPartialThinking(PartialThinking). The thinking is never part ofAiMessage.text(). Default:false.- Parameters:
returnThinking- whether to return the thinking- Returns:
- this builder
-
defaultRequestParameters
public JitLLMStreamingChatModel.JitLLMStreamingChatModelBuilder defaultRequestParameters(ChatRequestParameters defaultRequestParameters) Sets the parameters used for every request, unless the request sets them itself. Values set directly on this builder (for exampletemperature(Double)) take precedence.- Parameters:
defaultRequestParameters- the default request parameters- Returns:
- this builder
-
listeners
public JitLLMStreamingChatModel.JitLLMStreamingChatModelBuilder listeners(List<ChatModelListener> listeners) Sets the listeners notified about every request, response and error.- Parameters:
listeners- the listeners- Returns:
- this builder
-
build
Builds the model. This loads the model file, which can take a while for large models.- Returns:
- the model
-