generate(messages, **options) blocks and returns a GenerationResult. Input history is not mutated. generate_stream() returns a lazy, single-pass context-managed iterator:
# With the model still loaded:
withmodel.generate_stream(messages) asstream:
foreventinstream:
ifevent.type=="text":
print(event.text, end="", flush=True)
result=stream.result()
stream.result() drives any remaining work and returns the cached result. stream.cancel() stops cooperatively; exiting the context cleans up. Blocking calls accept a threading.Event as cancel_event. Cancellation returns partial work with stop_reason="cancelled"; failures raise GenerationError with partial_result and __cause__.
The result includes latest-round text, appendable new_messages, rounds, usage.input_tokens, usage.output_tokens, and duration_ms. Check stop_reason: complete, tool_calls, max_rounds, stop_condition, max_tokens, cancelled, context_limit, or stop_string. Latest-round text is not guaranteed to be a final answer.
Options include max_rounds (default 20), max_tokens_per_round, reasoning, stop_strings, and stop_when. Sampling overrides: temperature, top_p, top_k, min_p, repetition_penalty, presence_penalty, frequency_penalty, seed. Omitted sampling settings use model/runtime defaults. Reasoning counts toward the per-round token cap; stop strings apply to answer text and are removed from output.
on_text and on_reasoning are alternatives to reading those stream events. on_round_start identifies each round. Callbacks run inline; keep them short. Their exceptions propagate after cleanup.
All tools must provide executors or all must be manual. Executors receive validated arguments and return JSON-compatible values. An optional keyword-only context parameter receives ToolContext.cancel_event. Executors may run on background threads; async executors are unsupported.
Defaults: tool_execution="sequential", tool_execution_timing="immediate", tool_error_behavior="continue". Alternatives are "parallel" with optional max_concurrent_tools, "after_generation", and "stop", respectively. Executors are not automatically retried.
Notifications: on_tool_call, on_tool_start, on_tool_result, on_tool_error, on_tool_cancel, on_tool_validation_error. For manual handling, omit execute, append result.new_messages, execute result.tool_calls, and append tool_result(call, output) messages before generating again. Do not execute twice if you already handled an early notification.
Tools and structured output accept supported JSON Schema. The optional wfloat[schemas] extra also accepts Pydantic models and dataclasses, returning validated instances. Additional validators run after constrained decoding. Tools cannot be combined with structured output.
Correction attempts default to zero and consume the round budget. Exhausted validation raises; an expected early stop may return no validated output. output=None can represent either JSON null or absent output; inspect stop_reason.
Context and ownership
load_language_model(..., context_size=2048) sets capacity. model.context_size exposes it; model.count_input_tokens(messages, ...) counts the request, including matching tool/schema/reasoning options. Context-limit stops provide context_limit; Wfloat does not compact history automatically.
Same-model operations on different threads queue. Finish/close a stream before starting another operation on its driving thread. A managed executor cannot synchronously generate on its own active model.
Returns a handle immediately. result() resolves to the result; finished signals completion without returning it. cancel() stops cooperatively and resolves with stopReason: "cancelled". There is no pause/resume API. Actual failures reject with GenerationError, including partialResult and cause.
onText receives answer-text fragments; onReasoning receives reasoning separately. Both may span multiple rounds. onRoundStart({ roundIndex }) identifies a new round. Callbacks notify; they do not hold up inference or tool execution.
Option
Meaning
maxRounds
Maximum model rounds; default 20.
maxTokensPerRound
Output-token cap per round, including reasoning.
reasoning
Enable/disable reasoning where the model supports it.
stopStrings
Stop on matching answer text; omit the match from output.
stopWhen
After a round, inspect messages, rounds, and latestRound to prevent another round.
Sampling overrides; omitted settings use model/runtime defaults.
The result contains text from the latest round, newMessages to append to your history, rounds, stopReason, usage (inputTokens, outputTokens), and durationMs. Text is not necessarily a final answer: inspect the stop reason. Other reasons include complete, toolCalls, maxRounds, stopCondition, maxTokens, contextLimit, and stopString.
All tools must define execute, or all must omit it for manual handling. Executors receive validated arguments and { signal, callId, roundIndex }; return JSON-compatible values. JSON Schema and supported Zod schemas are accepted; install Zod yourself if using it.
Defaults: toolExecution: "sequential", toolExecutionTiming: "immediate", and toolErrorBehavior: "continue". Use "parallel" with optional maxConcurrentTools, "afterGeneration" to defer execution, or "stop" to fail on executor errors. Wfloat does not retry executors automatically.
Notifications: onToolCall, onToolStart, onToolResult, onToolError, onToolCancel, and onToolValidationError. Manual tools appear in result.toolCalls; append result.newMessages, then messages made with toolResult(call, output), before generating again. Calls observed early through onToolCall must not execute a second time when they appear in the result.
Constrains the answer to the supported JSON Schema subset and validates it, including supplied Zod validation. Tools cannot be combined with structured output. Correction attempts default to zero and consume the maxRounds budget. Exhausted validation rejects; an expected early stop can return without validated output. Stream consumers can reset their display at onRoundStart when corrections occur.
Context and ownership
loadLanguageModel(id, { contextSize }) sets capacity (default 2048). await model.countInputTokens(messages, options?) counts the formatted request; pass the same tools, reasoning, or structuredOutput settings. model.contextSize exposes capacity. Wfloat does not silently compact history; a context-limit stop includes contextLimit details.
Whole operations queue on one instance, including tool waits. An executor must not await another generation on the same model: that work would wait behind its own caller.
Returns a handle immediately. result() resolves to the result; finished signals completion without returning it. cancel() stops cooperatively and resolves with stopReason: "cancelled". There is no pause/resume API. Actual failures reject with GenerationError, including partialResult and cause.
onText receives answer-text fragments; onReasoning receives reasoning separately. Both may span multiple rounds. onRoundStart({ roundIndex }) identifies a new round. Callbacks notify; they do not hold up inference or tool execution.
Option
Meaning
maxRounds
Maximum model rounds; default 20.
maxTokensPerRound
Output-token cap per round, including reasoning.
reasoning
Enable/disable reasoning where the model supports it.
stopStrings
Stop on matching answer text; omit the match from output.
stopWhen
After a round, inspect messages, rounds, and latestRound to prevent another round.
Sampling overrides; omitted settings use model/runtime defaults.
The result contains text from the latest round, newMessages to append to your history, rounds, stopReason, usage (inputTokens, outputTokens), and durationMs. Text is not necessarily a final answer: inspect the stop reason. Other reasons include complete, toolCalls, maxRounds, stopCondition, maxTokens, contextLimit, and stopString.
All tools must define execute, or all must omit it for manual handling. Executors receive validated arguments and { signal, callId, roundIndex }; return JSON-compatible values. JSON Schema and supported Zod schemas are accepted; install Zod yourself if using it.
Defaults: toolExecution: "sequential", toolExecutionTiming: "immediate", and toolErrorBehavior: "continue". Use "parallel" with optional maxConcurrentTools, "afterGeneration" to defer execution, or "stop" to fail on executor errors. Wfloat does not retry executors automatically.
Notifications: onToolCall, onToolStart, onToolResult, onToolError, onToolCancel, and onToolValidationError. Manual tools appear in result.toolCalls; append result.newMessages, then messages made with toolResult(call, output), before generating again. Calls observed early through onToolCall must not execute a second time when they appear in the result.
Constrains the answer to the supported JSON Schema subset and validates it, including supplied Zod validation. Tools cannot be combined with structured output. Correction attempts default to zero and consume the maxRounds budget. Exhausted validation rejects; an expected early stop can return without validated output. Stream consumers can reset their display at onRoundStart when corrections occur.
Context and ownership
loadLanguageModel(id, { contextSize }) sets capacity (default 2048). await model.countInputTokens(messages, options?) counts the formatted request; pass the same tools, reasoning, or structuredOutput settings. model.contextSize exposes capacity. Wfloat does not silently compact history; a context-limit stop includes contextLimit details.
Whole operations queue on one instance, including tool waits. An executor must not await another generation on the same model: that work would wait behind its own caller.