llm_core 0.7.0 copy "llm_core: ^0.7.0" to clipboard
llm_core: ^0.7.0 copied to clipboard

Core abstractions for LLM (Large Language Model) interactions. Provides common interfaces, models, and utilities used by LLM backend implementations.

Changelog #

All notable changes to this project will be documented in this file.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

[Unreleased] #

0.7.0 - 2026-09-22 #

Added #

  • RetryUtil.retryingStream and ErrorHandlers.isRetryableStreamError: a streaming turn that fails before emitting anything is re-issued, whether the failure arrived as a non-2xx response or in-band as an error event. Once output has reached the caller the error is surfaced untouched, because re-running would deliver the turn's opening twice. A read timeout is deliberately excluded — it has already waited readTimeout, and retrying it reads as a hang.
  • LLMToolResult: a typed return for LLMTool.execute, carrying the text the model sees, an isError flag, metadata the model never sees, and contentParts for providers that accept more than text. execute still returns dynamic, so existing tools are unaffected.
  • LLMMessage.thinking, .toolName and .toolResult; LLMChunkMessage.toolName and .toolResult. toolName removes the toolCallId -> name join every backend and every consumer was re-implementing.
  • LLMUsage.cachedTokens and .cacheWriteTokens, both a subset of promptTokens.
  • LLMToolParam.minimum and .maximum, emitted in toJsonSchema() for integer and number.
  • LLMChatOptions.usagePerChunk: asks for token usage on every streamed chunk rather than only the final one. Honored by llm_vllm; ignored elsewhere. Deliberately excluded from the cache key, since it cannot change the generated text.
  • abortableStream, and abortTrigger on HttpClientHelper.sendStreamingRequest: cancelling a streamChat subscription now aborts the in-flight request instead of leaving it running.
  • MergedOptions.toChatOptions, which backends use to rebuild a tool round's options.

Fixed #

  • Validation.validateMessage documented a guard it could not enforce: LLMMessage(role: user, content: '') derives a single empty text part, so the empty-user-message check never fired for it. The behaviour is correct and deliberate — every backend accepts an empty user message except Anthropic, where ClaudeMessageConverter substitutes a placeholder — and is now documented and pinned by tests rather than left to be rediscovered.
  • An invalid tool call was answered with text a backend could not recognize as a failure, so llm_claude sent it to Anthropic without is_error and the model read the error message as data.
  • StreamToolExecutor dropped the model's reasoning when it assembled an assistant turn that called a tool, so it was unrecoverable from history.
  • Every backend rebuilt a tool round's options by hand and all five omitted useCache, cacheTtl and recordMetrics, so those settings stopped applying from the second round on.
  • ErrorHandlers.isRetryableError now rejects RequestAbortedException explicitly rather than by accident of string matching.

Changed #

  • LLMMessage.status is documented as strictly application-level and never serialized, with the exact key set toJson() emits pinned by a test. toolName replaces its former double duty as the tool-name channel; it is still populated and still read as a fallback.
  • LLMChunk.usage, .promptEvalCount and .evalCount may now appear on every chunk (see usagePerChunk); take the latest rather than accumulating.
  • All packages bumped to 0.7.0.

0.6.0 - 2026-09-17 #

Added #

  • LLMInvalidToolCall, LLMChunkMessage.invalidToolCalls and LLMResponse.invalidToolCalls: tool calls whose arguments don't decode are returned here, with the raw text and the parse error. This matches LangChain's invalid_tool_calls and the Vercel AI SDK.
  • LLMToolCall.argumentsError and LLMToolCall.partition.

Changed #

  • A length turn now returns its tool calls, split into valid and invalid, instead of dropping them. The finish reason stays length, as in the OpenAI API.
  • StreamToolExecutor runs only the valid calls. Each invalid call gets a tool error back (with a hint when the output token limit was hit) and is echoed with {} as its arguments.

0.5.0 - 2026-09-11 #

Added #

  • LLMFinishReason.resolve and LLMFinishReason.canBecomeToolCalls — the shared rule for classifying a tool-call turn, so every backend applies it identically instead of each re-deriving it.

Fixed #

  • chatResponse returned toolCalls: null for Ollama, Claude and Gemini. It captured calls only from chunks with done: true, and all three deliver complete calls on an earlier chunk. Calls are now attributed to the turn that carried them, and calls the tool loop already executed are still excluded.
  • LLMResponse.finishReason could contradict LLMResponse.toolCalls. The finish reason is now resolved against the calls actually returned, so a backend that does not classify the turn itself still yields a consistent response.

Changed #

  • A turn that ends while carrying complete tool calls now reports LLMFinishReason.toolCalls, whatever the provider spelled. The OpenAI specification defines finish_reason as tool_calls "if the model called a tool", and providers violate it routinely. Code branching on LLMFinishReason.stop for a tool-calling turn must move to toolCalls. length, contentFilter and refusal are never reclassified: a truncated call is not executable, and a declined turn must stay visibly declined.

0.4.0 - 2026-08-30 #

Added #

  • LLMToolCallDelta and LLMChunkMessage.toolCallDeltas — fragments of a tool call that is still streaming. Backends announce the tool's name in their first event, so a caller can show which tool is running without waiting for its arguments.

Fixed #

  • Tool calls that take no arguments always failed. A zero-parameter tool is routinely called with "" (OpenAI-compatible servers) or with nothing to concatenate (Anthropic), and decoding that as JSON threw — the executor answered Tool x failed: FormatException and the message converters that replay history threw outright. LLMToolCall.argumentsJson now reads no arguments as an empty map; genuinely malformed JSON still throws.
  • Token counts survive a turn that ends in more than one done chunk. Several backends report the finish reason first and token usage in a trailing frame, and a tool loop produces a done chunk per round; counts were assigned unconditionally, so a later count-less chunk erased them.
  • 529 is retryable by default. It is Anthropic's transient "overloaded" signal, and its absence meant those failures were never retried.
  • ToolLoopIncompleteException.attemptsUsed reported 0 no matter how many tool rounds had run. Each recursion builds a fresh executor whose budget is already decremented, so the frame that runs out cannot see the rounds behind it; the accounting is now restated as the error unwinds, and the outermost frame — which knows the real ceiling — supplies the final number.

Changed #

  • LLMChunkMessage.toolCalls is documented as only ever holding complete, executable calls. Deltas are never executable and never appear there.

0.3.2 - 2026-08-18 #

Changed #

  • Version bumped to 0.3.2 in lockstep with the other packages. No changes to this package.

0.3.1 - 2026-08-18 #

Added #

  • createLLMHttpClient() — the default HTTP client for every backend. Applies TimeoutConfig.connectionTimeout, bounds the pool per host, and retires idle connections after 3s.
  • WriteGatedHttpClient — bounds how many requests may be connecting and writing at once (4 slots on macOS/iOS) to work around a Dart VM kqueue defect that loses writable events. See docs/concurrent-send-stall.md.
  • LLMChunkMessage.rawContent — the assistant turn as the model emitted it, tool-call markup intact. Set by local-inference backends.
  • RetryUtil.executeWithRetry accepts an onRetry callback and warns on every retry.

Changed #

  • HttpClientHelper.sendStreamingRequest now applies a timeout to send() by default. It previously did not, so a request that wedged before response headers arrived never recovered.
  • Streaming requests are built with http.Request + bodyBytes instead of StreamedRequest.
  • Dependency floors raised: Dart SDK ^3.12.0 (was ^3.8.0), http ^1.6.0, lints ^6.1.0, test ^1.31.0.

0.3.0 - 2026-08-17 #

Added #

  • ReasoningEffort enum (none…max) and LLMChatOptions.reasoningEffort — a portable reasoning-depth knob alongside reasoningBudget.
  • reasoningEffortForBudget() for backends without a native token budget.
  • LLMUsage.reasoningTokens for providers that report reasoning-token usage.

0.2.0 - 2026-08-12 #

Fixed #

  • LLMFinishReason.fromProvider now recognizes stop_sequence, model_context_window_exceeded, refusal, and Gemini's RECITATION / PROHIBITED_CONTENT / BLOCKLIST / SPII / IMAGE_SAFETY, which previously all collapsed to unknown.

Added #

  • LLMFinishReason.refusal for provider safety declines, which arrive as successful responses with empty or partial content.
  • Typed message content parts, typed tool calls, provider capabilities, response usage, finish reasons, thinking output, and provider metadata.
  • LLMChatOptions for generation, reasoning, tool behavior, structured output, timeout, retry, cache, metrics, and backend-specific options.
  • Shared repository feature helpers for cache and metrics handling.
  • LLMResponseFormat sealed class hierarchy for structured output:
    • JsonFormat — simple JSON mode; instructs the model to produce valid JSON without schema enforcement
    • JsonSchemaFormat({required name, required schema, strict = true}) — full JSON Schema mode; schema is forwarded to the backend verbatim
    • Both are const-constructible and work with exhaustive switch pattern matching
  • responseFormat field on StreamChatOptions (nullable, defaults to null; fully backward compatible)
  • StreamChatOptionsMerger and MergedOptions now carry and propagate responseFormat

Changed #

  • Breaking: Core chat APIs now accept LLMChatOptions?; StreamChatOptions remains as a compatibility alias.
  • Cache keys now use stable JSON-shaped request data instead of object stringification.

0.1.9 - 2026-02-28 #

Changed #

  • Breaking: Removed requireFinalAssistantResponse option from StreamChatOptions, StreamChatOptionsMerger, and StreamToolExecutor. Tool loops now always require a final assistant response — there is no reason to allow tool loops to end without the assistant reporting back.
  • maxToolAttempts default increased from 25 to 90 across all repositories and builders.
  • chatResponse() tool loop detection refined: only actual tool result chunks (LLMRole.tool) trigger the incomplete-loop check, not tool calls appearing alongside content.

0.1.8 - 2026-02-26 #

Added #

  • Tool result chunks emitted to stream: StreamToolExecutor now yields LLMChunk with role: LLMRole.tool after each tool execution, so chat consumers can display "Tool X returned: Y" per OpenAI function calling specs
  • LLMToolCall.toApiFormat() helper for converting to OpenAI/Ollama API format
  • Assistant message with tool_calls added to message history before tool results (API-compliant sequence)
  • Content accumulation from stream chunks for assistant messages that include both text and tool calls

Changed #

  • Breaking: Removed toolName from LLMMessage and LLMChunkMessage; use toolCallId only (OpenAI canonical format)
  • Breaking: Tool message validation now requires toolCallId (removed toolName option)
  • StreamToolExecutor accumulates content and thinking from chunks for the assistant message

0.1.7 - 2026-02-10 #

Added #

  • batchEmbed() on LLMChatRepository: explicit API for embedding multiple texts in one call. Same signature as embed(); default implementation delegates to embed(). Documented for Ollama, OpenAI, and llama.cpp backends.

0.1.6 - 2026-02-10 #

Fixed #

  • Hardened StreamToolExecutor to always synthesize a non-empty toolCallId for LLMRole.tool messages when a backend-provided LLMToolCall.id is missing or empty, preventing Tool message must have toolCallId validation errors.
  • Improved tool execution error handling so that thrown tool exceptions are surfaced as tool messages rather than crashing the stream.

0.1.5 - 2026-01-26 #

Added #

  • StreamChatOptions class to encapsulate all streaming chat options and reduce parameter proliferation
  • RetryConfig and RetryUtil for configurable retry logic with exponential backoff
  • TimeoutConfig for flexible timeout configuration (connection, read, total, large payloads)
  • LLMMetrics interface and DefaultLLMMetrics implementation for optional metrics collection
  • chatResponse() method on LLMChatRepository for non-streaming complete responses
  • Input validation utilities in Validation class
  • ChatRepositoryBuilderBase for implementing builder patterns in repository implementations
  • StreamChatOptionsMerger for merging options from multiple sources
  • HTTP client utilities (HttpClientHelper) for consistent request handling
  • Error handling utilities (ErrorHandlers, BackendErrorHandler) for standardized error processing
  • Tool execution utilities (ToolExecutor) for managing tool calling workflows

Changed #

  • streamChat() now accepts optional StreamChatOptions parameter
  • Improved error handling and retry logic across all backends
  • Enhanced documentation

0.1.0 - 2026-01-19 #

Added #

  • Initial release
  • Core abstractions for LLM interactions:
    • LLMChatRepository - Abstract interface for chat completions
    • LLMMessage - Message representation with roles and content
    • LLMResponse - Response wrapper with metadata
    • LLMChunk - Streaming response chunks
    • LLMEmbedding - Text embedding representation
  • Tool calling support:
    • LLMTool - Tool definition with JSON Schema parameters
    • LLMToolCall - Tool invocation representation
    • LLMToolParam - Parameter definitions
  • Exception types for error handling
0
likes
160
points
1.06k
downloads

Documentation

API reference

Publisher

unverified uploader

Weekly Downloads

Core abstractions for LLM (Large Language Model) interactions. Provides common interfaces, models, and utilities used by LLM backend implementations.

Repository (GitHub)
View/report issues
Contributing

Topics

#llm #ai #chat #embeddings #tools

License

MIT (license)

Dependencies

http, logging

More

Packages that depend on llm_core