- Simplified token type assignment in OAuth service. - Removed unused session_lock module and re-exported Session from zesdex_entities. - Cleaned up session entity by removing unnecessary comments and code. - Consolidated session handling in HTTP handlers for better readability. - Improved formatting and readability in OAuth repository tests. - Enhanced session lock repository with clearer match statements. - Streamlined session repository error handling. - Refined RNG tests for better clarity. - Adjusted module visibility and organization in lib.rs. - Updated IPC client and connection code for better error handling and clarity. - Improved frame handling in IPC for better readability. - Organized module imports and added test utilities for IPC. - Enhanced database connection error handling. - Simplified JWT token creation error handling. - Improved password verification error handling. - Cleaned up state management code for better readability. - Refactored middleware for session authentication and rate limiting. - Simplified clipboard utility for better error handling. - Enhanced logging initialization for better error reporting. - Improved pagination utility with clearer method annotations. - Cleaned up sanitization functions for filenames and paths. - Enhanced slug generation functions for better clarity and usability.
68 lines
2.4 KiB
Rust
68 lines
2.4 KiB
Rust
//! Unified token-count estimation for context-window budgeting.
|
|
//!
|
|
//! Flow: text -> `tiktoken_rs::o200k_base_singleton()` (BPE vocab embedded
|
|
//! in the binary via `include_str!`, no network access) -> `encode_ordinary`
|
|
//! -> token count.
|
|
//!
|
|
//! Why: replaces three independent char-count heuristics that disagreed
|
|
//! with each other (`/3` in the old `shortsend.rs`, `/4` in the turn
|
|
//! loop, `/4` again in the status bar) with one real BPE tokenizer.
|
|
//! `o200k_base` is an approximation for non-OpenAI providers but is far
|
|
//! closer than a flat byte-per-token guess; it's only used for the
|
|
//! 85%/95% budget thresholds, not for billing-accurate counts.
|
|
|
|
|
|
/// Count tokens in a single string under `o200k_base`.
|
|
///
|
|
/// Return: the BPE token count for `text`. `encode_ordinary` (not
|
|
/// `encode`/`encode_with_special_tokens`) is used deliberately — message
|
|
/// content that happens to contain a special-token-shaped substring
|
|
/// (e.g. literal text `<|endoftext|>` pasted by a user) must be counted
|
|
/// as ordinary text, not interpreted as a control token.
|
|
pub fn count_tokens(text: &str) -> usize {
|
|
tiktoken_rs::o200k_base_singleton()
|
|
.encode_ordinary(text)
|
|
.len()
|
|
}
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
use crate::dto::chat::message::ChatMessage;
|
|
|
|
/// Count tokens in a `ChatMessage`'s text content.
|
|
fn count_message_tokens(msg: &ChatMessage) -> usize {
|
|
msg.content.as_deref().map_or(0, count_tokens)
|
|
}
|
|
|
|
#[test]
|
|
fn empty_string_has_zero_tokens() {
|
|
assert_eq!(count_tokens(""), 0);
|
|
}
|
|
|
|
#[test]
|
|
fn known_short_phrase_has_expected_token_count() {
|
|
// Verified empirically against tiktoken-rs 0.12's o200k_base:
|
|
// "hello world" -> [24912, 2375], i.e. 2 tokens.
|
|
assert_eq!(count_tokens("hello world"), 2);
|
|
}
|
|
|
|
#[test]
|
|
fn known_code_snippet_has_expected_token_count() {
|
|
// Verified empirically: 9 tokens under o200k_base.
|
|
assert_eq!(count_tokens("fn main() { println!(\"hi\"); }"), 9);
|
|
}
|
|
|
|
#[test]
|
|
fn message_with_no_content_counts_zero() {
|
|
let msg = ChatMessage::assistant(None);
|
|
assert_eq!(count_message_tokens(&msg), 0);
|
|
}
|
|
|
|
#[test]
|
|
fn message_token_count_matches_count_tokens_on_its_content() {
|
|
let msg = ChatMessage::user("hello world");
|
|
assert_eq!(count_message_tokens(&msg), count_tokens("hello world"));
|
|
}
|
|
}
|