llm_core 0.5.0 copy "llm_core: ^0.5.0" to clipboard
llm_core: ^0.5.0 copied to clipboard

Core abstractions for LLM (Large Language Model) interactions. Provides common interfaces, models, and utilities used by LLM backend implementations.

Changelog #

All notable changes to this project will be documented in this file.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

[Unreleased] #

0.5.0 - 2026-09-11 #

Added #

  • LLMFinishReason.resolve and LLMFinishReason.canBecomeToolCalls — the shared rule for classifying a tool-call turn, so every backend applies it identically instead of each re-deriving it.

Fixed #

  • chatResponse returned toolCalls: null for Ollama, Claude and Gemini. It captured calls only from chunks with done: true, and all three deliver complete calls on an earlier chunk. Calls are now attributed to the turn that carried them, and calls the tool loop already executed are still excluded.
  • LLMResponse.finishReason could contradict LLMResponse.toolCalls. The finish reason is now resolved against the calls actually returned, so a backend that does not classify the turn itself still yields a consistent response.

Changed #

  • A turn that ends while carrying complete tool calls now reports LLMFinishReason.toolCalls, whatever the provider spelled. The OpenAI specification defines finish_reason as tool_calls "if the model called a tool", and providers violate it routinely. Code branching on LLMFinishReason.stop for a tool-calling turn must move to toolCalls. length, contentFilter and refusal are never reclassified: a truncated call is not executable, and a declined turn must stay visibly declined.

0.4.0 - 2026-08-30 #

Added #

  • LLMToolCallDelta and LLMChunkMessage.toolCallDeltas — fragments of a tool call that is still streaming. Backends announce the tool's name in their first event, so a caller can show which tool is running without waiting for its arguments.

Fixed #

  • Tool calls that take no arguments always failed. A zero-parameter tool is routinely called with "" (OpenAI-compatible servers) or with nothing to concatenate (Anthropic), and decoding that as JSON threw — the executor answered Tool x failed: FormatException and the message converters that replay history threw outright. LLMToolCall.argumentsJson now reads no arguments as an empty map; genuinely malformed JSON still throws.
  • Token counts survive a turn that ends in more than one done chunk. Several backends report the finish reason first and token usage in a trailing frame, and a tool loop produces a done chunk per round; counts were assigned unconditionally, so a later count-less chunk erased them.
  • 529 is retryable by default. It is Anthropic's transient "overloaded" signal, and its absence meant those failures were never retried.
  • ToolLoopIncompleteException.attemptsUsed reported 0 no matter how many tool rounds had run. Each recursion builds a fresh executor whose budget is already decremented, so the frame that runs out cannot see the rounds behind it; the accounting is now restated as the error unwinds, and the outermost frame — which knows the real ceiling — supplies the final number.

Changed #

  • LLMChunkMessage.toolCalls is documented as only ever holding complete, executable calls. Deltas are never executable and never appear there.

0.3.2 - 2026-08-18 #

Changed #

  • Version bumped to 0.3.2 in lockstep with the other packages. No changes to this package.

0.3.1 - 2026-08-18 #

Added #

  • createLLMHttpClient() — the default HTTP client for every backend. Applies TimeoutConfig.connectionTimeout, bounds the pool per host, and retires idle connections after 3s.
  • WriteGatedHttpClient — bounds how many requests may be connecting and writing at once (4 slots on macOS/iOS) to work around a Dart VM kqueue defect that loses writable events. See docs/concurrent-send-stall.md.
  • LLMChunkMessage.rawContent — the assistant turn as the model emitted it, tool-call markup intact. Set by local-inference backends.
  • RetryUtil.executeWithRetry accepts an onRetry callback and warns on every retry.

Changed #

  • HttpClientHelper.sendStreamingRequest now applies a timeout to send() by default. It previously did not, so a request that wedged before response headers arrived never recovered.
  • Streaming requests are built with http.Request + bodyBytes instead of StreamedRequest.
  • Dependency floors raised: Dart SDK ^3.12.0 (was ^3.8.0), http ^1.6.0, lints ^6.1.0, test ^1.31.0.

0.3.0 - 2026-08-17 #

Added #

  • ReasoningEffort enum (nonemax) and LLMChatOptions.reasoningEffort — a portable reasoning-depth knob alongside reasoningBudget.
  • reasoningEffortForBudget() for backends without a native token budget.
  • LLMUsage.reasoningTokens for providers that report reasoning-token usage.

0.2.0 - 2026-08-12 #

Fixed #

  • LLMFinishReason.fromProvider now recognizes stop_sequence, model_context_window_exceeded, refusal, and Gemini's RECITATION / PROHIBITED_CONTENT / BLOCKLIST / SPII / IMAGE_SAFETY, which previously all collapsed to unknown.

Added #

  • LLMFinishReason.refusal for provider safety declines, which arrive as successful responses with empty or partial content.
  • Typed message content parts, typed tool calls, provider capabilities, response usage, finish reasons, thinking output, and provider metadata.
  • LLMChatOptions for generation, reasoning, tool behavior, structured output, timeout, retry, cache, metrics, and backend-specific options.
  • Shared repository feature helpers for cache and metrics handling.
  • LLMResponseFormat sealed class hierarchy for structured output:
    • JsonFormat — simple JSON mode; instructs the model to produce valid JSON without schema enforcement
    • JsonSchemaFormat({required name, required schema, strict = true}) — full JSON Schema mode; schema is forwarded to the backend verbatim
    • Both are const-constructible and work with exhaustive switch pattern matching
  • responseFormat field on StreamChatOptions (nullable, defaults to null; fully backward compatible)
  • StreamChatOptionsMerger and MergedOptions now carry and propagate responseFormat

Changed #

  • Breaking: Core chat APIs now accept LLMChatOptions?; StreamChatOptions remains as a compatibility alias.
  • Cache keys now use stable JSON-shaped request data instead of object stringification.

0.1.9 - 2026-02-28 #

Changed #

  • Breaking: Removed requireFinalAssistantResponse option from StreamChatOptions, StreamChatOptionsMerger, and StreamToolExecutor. Tool loops now always require a final assistant response — there is no reason to allow tool loops to end without the assistant reporting back.
  • maxToolAttempts default increased from 25 to 90 across all repositories and builders.
  • chatResponse() tool loop detection refined: only actual tool result chunks (LLMRole.tool) trigger the incomplete-loop check, not tool calls appearing alongside content.

0.1.8 - 2026-02-26 #

Added #

  • Tool result chunks emitted to stream: StreamToolExecutor now yields LLMChunk with role: LLMRole.tool after each tool execution, so chat consumers can display "Tool X returned: Y" per OpenAI function calling specs
  • LLMToolCall.toApiFormat() helper for converting to OpenAI/Ollama API format
  • Assistant message with tool_calls added to message history before tool results (API-compliant sequence)
  • Content accumulation from stream chunks for assistant messages that include both text and tool calls

Changed #

  • Breaking: Removed toolName from LLMMessage and LLMChunkMessage; use toolCallId only (OpenAI canonical format)
  • Breaking: Tool message validation now requires toolCallId (removed toolName option)
  • StreamToolExecutor accumulates content and thinking from chunks for the assistant message

0.1.7 - 2026-02-10 #

Added #

  • batchEmbed() on LLMChatRepository: explicit API for embedding multiple texts in one call. Same signature as embed(); default implementation delegates to embed(). Documented for Ollama, OpenAI, and llama.cpp backends.

0.1.6 - 2026-02-10 #

Fixed #

  • Hardened StreamToolExecutor to always synthesize a non-empty toolCallId for LLMRole.tool messages when a backend-provided LLMToolCall.id is missing or empty, preventing Tool message must have toolCallId validation errors.
  • Improved tool execution error handling so that thrown tool exceptions are surfaced as tool messages rather than crashing the stream.

0.1.5 - 2026-01-26 #

Added #

  • StreamChatOptions class to encapsulate all streaming chat options and reduce parameter proliferation
  • RetryConfig and RetryUtil for configurable retry logic with exponential backoff
  • TimeoutConfig for flexible timeout configuration (connection, read, total, large payloads)
  • LLMMetrics interface and DefaultLLMMetrics implementation for optional metrics collection
  • chatResponse() method on LLMChatRepository for non-streaming complete responses
  • Input validation utilities in Validation class
  • ChatRepositoryBuilderBase for implementing builder patterns in repository implementations
  • StreamChatOptionsMerger for merging options from multiple sources
  • HTTP client utilities (HttpClientHelper) for consistent request handling
  • Error handling utilities (ErrorHandlers, BackendErrorHandler) for standardized error processing
  • Tool execution utilities (ToolExecutor) for managing tool calling workflows

Changed #

  • streamChat() now accepts optional StreamChatOptions parameter
  • Improved error handling and retry logic across all backends
  • Enhanced documentation

0.1.0 - 2026-01-19 #

Added #

  • Initial release
  • Core abstractions for LLM interactions:
    • LLMChatRepository - Abstract interface for chat completions
    • LLMMessage - Message representation with roles and content
    • LLMResponse - Response wrapper with metadata
    • LLMChunk - Streaming response chunks
    • LLMEmbedding - Text embedding representation
  • Tool calling support:
    • LLMTool - Tool definition with JSON Schema parameters
    • LLMToolCall - Tool invocation representation
    • LLMToolParam - Parameter definitions
  • Exception types for error handling
0
likes
160
points
1.05k
downloads

Documentation

API reference

Publisher

unverified uploader

Weekly Downloads

Core abstractions for LLM (Large Language Model) interactions. Provides common interfaces, models, and utilities used by LLM backend implementations.

Repository (GitHub)
View/report issues
Contributing

Topics

#llm #ai #chat #embeddings #tools

License

MIT (license)

Dependencies

http, logging

More

Packages that depend on llm_core