@azghr/shorn
shorn
Truncate strings by byte budget without ever breaking a grapheme cluster. Safe for database columns, HTTP headers, and anything with a hard byte limit.
The problem
A MySQL VARCHAR(255) utf8mb4 column holds up to 255 UTF-8 bytes. If you str.slice(0, 255) a user bio, you can land in the middle of a 4-byte emoji or a ZWJ family sequence. Encoding the sliced string back to UTF-8 and decoding it produces U+FFFD (replacement character) — the emoji is destroyed.
The same failure mode hits HTTP headers, MongoDB indexed fields, and any system with a byte limit on encoded text. The fix is to never split a grapheme cluster, which is what shorn guarantees.
Install
npm install @azghr/shorn
# or
pnpm add @azghr/shorn
# or
yarn add @azghr/shorn
Use
Truncate a string so it fits in 64 UTF-8 bytes. If it already fits, the original string reference is returned unchanged (zero alloc).
import { shorn } from "@azghr/shorn";
const bio = "Hi! I love 🎉 and 👨👩👧👦";
const safe = shorn(bio, 64);
// safe is 42 bytes, contains whole graphemes only
Truncate with an ellipsis counted inside the budget:
const longBio = "Hi! I love 🎉 and 👨👩👧👦 and also 🌍 café 🇩🇪";
const safe = shorn(longBio, 50, { unit: "utf8", ellipsis: "..." });
// safe is at most 50 bytes, ends with "..."
Measure in UTF-16 code units or grapheme clusters instead of bytes:
shorn(bio, 30, { unit: "utf16" }); // 30 code units
shorn(bio, 10, { unit: "grapheme" }); // 10 grapheme clusters
shorn(bio, 10, { unit: "grapheme", ellipsis: "…" }); // 10 grapheme clusters + ellipsis
API
shorn(input, max, options?)
| Param | Type | Default | Description |
|---|---|---|---|
input |
string |
(required) | String to truncate |
max |
number |
(required) | Maximum size in the chosen unit. Must be a non-negative integer |
options |
ShornOptions |
{} |
See below |
Returns: string -- the longest prefix of whole grapheme clusters whose measured size does not exceed max, optionally followed by the ellipsis. If the input already fits, returns the same reference (identity fast path).
Throws: TypeError if max < 0, NaN, Infinity, or not an integer.
ShornOptions
interface ShornOptions {
unit?: "utf8" | "utf16" | "grapheme"; // default "utf8"
ellipsis?: string; // default "" (none)
locale?: string; // passed to Intl.Segmenter
}
unit: Measurement mode."utf8"usesTextEncoderbyte length."utf16"usesstring.length(code units)."grapheme"usesIntl.Segmenterwithgranularity: "grapheme".ellipsis: Appended to the result when truncation occurs. Measured in the sameunitand counted inside the budget. If not even one grapheme can fit alongside the ellipsis, the ellipsis is dropped and only the grapheme prefix is returned. If the ellipsis alone exceeds the budget, it is dropped.locale: BCP-47 locale forwarded toIntl.Segmenter. Only affects grapheme segmentation rules (e.g. locale-specific cluster boundaries in Indic scripts). Defaultundefined(runtime default).
Non-goals
This package will never support:
- Streaming / chunked truncation
- Unicode normalization (NFC/NFD/NFKC/NFKD)
- Custom segmenter rules -- only
Intl.Segmenterwith"grapheme"granularity - Word-boundary or sentence-boundary truncation
- In-place mutation
- Non-UTF encodings (GB18030, Shift_JIS, etc.)
- Surrogate-pair healing (grapheme segmentation already handles this)
- Configurable normalization of the ellipsis
- Terminal column width measurement
- Padding
Composition example: If you need word-boundary truncation, pair shorn with a word splitter:
import { shorn } from "@azghr/shorn";
function truncateWords(text: string, maxBytes: number): string {
const words = text.split(" ");
let candidate = "";
for (const word of words) {
const next = candidate ? candidate + " " + word : word;
if (shorn(next, maxBytes).length !== next.length) break;
candidate = next;
}
return candidate;
}
Related Packages
- @azghr/filterkit — Framework-agnostic, type-safe filtering for TypeScript
- @azghr/singlet — Deduplicate concurrent async calls
- forbear — Read server rate-limit instructions from HTTP responses
- quiesce — Ordered, timeboxed graceful shutdown for Node
- sortition — Deterministic percentage rollouts and A/B bucketing
License
MIT