Skip to main content

Press o200k_base or cl100k_base to count text tokens, and check the Unicode code point, UTF-8 byte count and Token ID at the same time. Suitable for comparing text lengths without passing the results off as model bills.

Statistics raw text by selected encoding, not a universal result for all models, nor a bill. Chat wrappers, tool definitions, images, and hidden prompts do not count. The word segmentation table may be a little slow to load for the first time.

20 / 20,000 characters

Shortcut key: Ctrl/⌘+Enter. Modifying inputs or options clears old results.

Instructions for use

  1. Select the corresponding word segmentation encoding according to the target model data.
  2. Paste the text or load the mixed Chinese and English code examples, and click Statistics Token.
  3. Check the counting card and Token ID; recalculate after switching encoding.

Input and output examples

Example input
hello(cl100k_base)
Example output
Number of tokens: 1 Unicode code point: 5 UTF-8 bytes: 5

The word segmentation boundaries of different texts and different encodings may be different, and a fixed word count ratio cannot be used to replace the actual word segmentation.

FAQ

Only the original text entered is counted here, and chat packaging, tool definitions, pictures, hidden prompts or output content are not included. Actual billing is based on the usage returned by the service provider.

No. Language, spacing, punctuation, and code all affect word segmentation. When the ID list exceeds 200, only the first 200 will be displayed, but the total number of Tokens will still be counted and entered completely.

Calculation basis and reference materials

Use gpt-tokenizer's o200k_base or cl100k_base encoding to tokenize plain text, displaying both the number of code points and the number of UTF-8 bytes. Does not include additional usage for chat packages, images or tool calls.

Content check: · About this site and content description

Related tools