Results 13 packages
Sort by

dart_sentencepiece_tokenizercopy "dart_sentencepiece_tokenizer: ^1.4.1" to clipboard
dart_sentencepiece_tokenizer: ^1.4.1 copied to clipboard

7
likes
160
points
36k
downloads
A lightweight pure Dart SentencePiece tokenizer supporting BPE (Gemma), Unigram (Llama), and Hugging Face tokenizer.json pipelines.#nlp#sentencepiece#tokenizer#machine-learning#llm

val_highlightcopy "val_highlight: ^0.1.0" to clipboard
val_highlight: ^0.1.0 copied to clipboard

recently created package Added 9 days ago
6
likes
160
points
5.2k
downloads
Syntax highlighting for 55 languages in pure Dart. Incremental updates, themes, VS Code theme import, and HTML output. Runs on every platform.#syntax-highlighting#highlighter#code#tokenizer#html

tiktoken_tokenizer_gpt4o_o1copy "tiktoken_tokenizer_gpt4o_o1: ^1.2.1" to clipboard
tiktoken_tokenizer_gpt4o_o1: ^1.2.1 copied to clipboard

3
likes
140
points
867
downloads
OpenAI's Tiktoken tokenizer for models: GPT-4, GPT-4o, GPT-4o-mini, o1, o1-mini, and o1-preview.#tokenizer#openai#gpt#gpt-4o#tiktoken

dart_bert_tokenizercopy "dart_bert_tokenizer: ^1.3.0" to clipboard
dart_bert_tokenizer: ^1.3.0 copied to clipboard

2
likes
160
points
717
downloads
A pure Dart BERT WordPiece tokenizer with Hugging Face tokenizer.json loading, pre-tokenized alignment, dynamic AddedTokens, Unicode offsets, and verified model fixtures.#nlp#bert#tokenizer#machine-learning#wordpiece

tiny_segmenter_dartcopy "tiny_segmenter_dart: ^1.0.1" to clipboard
tiny_segmenter_dart: ^1.0.1 copied to clipboard

4
likes
135
points
201
downloads
A compact Japanese text tokenizer for Dart. TinySegmenter is a Japanese word segmentation library based on the original JavaScript implementation by Taku Kudo.#japanese#text-processing#tokenizer#nlp#segmentation

hf_tokenizerscopy "hf_tokenizers: ^1.2.2" to clipboard
hf_tokenizers: ^1.2.2 copied to clipboard

1
likes
160
points
679
downloads
HuggingFace tokenizers for Dart over FFI. Load any tokenizer.json and get byte-exact BPE, WordPiece, and Unigram encoding, backed by the Rust crate.#llm#ai#tokenizer#ffi#nlp
screenshot

model2veccopy "model2vec: ^2.0.3" to clipboard
model2vec: ^2.0.3 copied to clipboard

2
likes
150
points
80
downloads
On-device Model2Vec text embeddings for Dart & Flutter — a self-contained Rust core via FFI and Native Assets. Fast, local, static, minimal memory.#rag#nlp#embeddings#tokenizer#model2vec

betto_icucopy "betto_icu: ^0.1.0" to clipboard
betto_icu: ^0.1.0 copied to clipboard

0
likes
160
points
380
downloads
Unicode text tokenization for Dart — Tokenizer interface, IcuTokenizer (system ICU FFI, UAX #29), and RegExpTokenizer (pure Dart, Latin fallback). #text#unicode#nlp#tokenizer#icu

flutter_syntax_highlightcopy "flutter_syntax_highlight: ^0.2.0" to clipboard
flutter_syntax_highlight: ^0.2.0 copied to clipboard

0
likes
160
points
244
downloads
Dart-only syntax highlighting in two layers: a pure-Dart tokenizer underneath and a thin widget on top. The tokens rejoin into the exact input, byte for byte.#syntax-highlighting#tokenizer#code#text#widget
screenshot

sqlite3_jiebacopy "sqlite3_jieba: ^0.2.1" to clipboard
sqlite3_jieba: ^0.2.1 copied to clipboard

recently created package Added 26 days ago
0
likes
150
points
320
downloads
A jieba-backed FTS5 tokenizer for SQLite, shipped as a native asset. Adds a `jieba` tokenizer and a `jieba_cut()` function to package:sqlite3.#sqlite#fts5#full-text-search#chinese#tokenizer

Check our help page for details on search expressions and result ranking.