interface LlamaCppInputsLlamaBaseCppInputsBaseLLMParamsPrompt processing batch size.
Text context size.
Embedding mode only.
Use fp16 for KV cache.
GBNF string to be used to format output. Also known as grammar.
Number of layers to store in VRAM.
JSON schema to be used to format output. Also known as grammar.
The llama_eval() call computes all logits, not just the last one.
The maximum number of concurrent calls that can be made.
Defaults to Infinity, which means no limit.
The maximum number of retries that can be made for a single call, with an exponential backoff between each attempt. Defaults to 6.
Path to the model on the filesystem.
Custom handler to handle failed attempts. Takes the originally thrown error object as input, and should itself throw an error if the input error is not retryable.
Add the begining of sentence token.
If null, a random seed will be used.
The randomness of the responses, e.g. 0.1 deterministic, 1.5 creative, 0.8 balanced, 0 disables.
Number of threads to use to evaluate tokens.
Consider the n most likely tokens, where n is 1 to vocabulary size, 0 disables (uses full vocabulary). Note: only applies when temperature > 0.
Selects the smallest token set whose probability exceeds P, where P is between 0 - 1, 1 disables. Note: only applies when temperature > 0.
Trim whitespace from the end of the generated text Disabled by default.
Force system to keep model in RAM.
Use mmap if possible.
Only load the vocabulary, no weights.
Note that the modelPath is the only required parameter. For testing you
can set this in the environment variable LLAMA_PATH.