Introduction
Strings power almost every user-facing feature: search, display, validation, formatting, URLs, filenames, logs, and more. Yet many teams repeatedly re-implement the same helpers across projects—or reach for heavyweight libraries for small jobs.
This guide gives you 15 practical, copy-paste-ready JavaScript string manipulation functions. They’re small, focused, and designed for everyday use in both Node.js and the browser. Each comes with examples and notes on edge cases so you can drop them into your codebase with confidence.
What you’ll get:
- Focused, pure utility functions that don’t mutate their inputs
- Unicode-aware approaches where it matters (diacritics, graphemes)
- Practical examples and real-world recipes
Let’s level up your string toolbox.
How to use this guide
- Each utility stands alone—copy only what you need.
- When applicable, functions include options for locale, safety, or performance trade-offs.
- Examples demonstrate common pitfalls and best practices.
1) capitalize: Capitalize the first character (optionally lower the rest)
Useful for names, titles, and labels. Handles the first character robustly and optionally normalizes the remaining string.
/**
* Capitalize the first character of a string.
* Optionally lowercases the rest. Attempts to honor locale when provided.
*
* @param {string} str
* @param {Object} [opts]
* @param {boolean} [opts.lowerRest=false]
* @param {string} [opts.locale] e.g. 'en-US', 'tr', etc.
* @returns {string}
*/
function capitalize(str, opts = {}) {
const { lowerRest = false, locale } = opts;
if (typeof str !== 'string' || str.length === 0) return '';
// Try to segment the first grapheme (handles emoji, compound characters)
let first = '';
let rest = '';
if (typeof Intl !== 'undefined' && Intl.Segmenter) {
const seg = new Intl.Segmenter(locale, { granularity: 'grapheme' }).segment(str);
const it = seg[Symbol.iterator]();
const { value } = it.next();
first = value.segment;
rest = str.slice(value.index + value.segment.length);
} else {
// Fallback: not perfect for all scripts, but works for most
const chars = Array.from(str);
first = chars[0];
rest = chars.slice(1).join('');
}
const firstCap = locale ? first.toLocaleUpperCase(locale) : first.toUpperCase();
const tail = lowerRest ? (locale ? rest.toLocaleLowerCase(locale) : rest.toLowerCase()) : rest;
return firstCap + tail;
}
// Examples
// capitalize('hello world') -> 'Hello world'
// capitalize('ßtraße', { lowerRest: true, locale: 'de' }) -> 'Sstraße'
Actionable advice:
- Use lowerRest: true when normalizing labels or titles for consistent casing.
- Provide a locale for specific languages like Turkish (“i/I”).
2) toTitleCase: Human-friendly title casing (with small word exclusions)
Simple title-casing that keeps “small words” lowercase unless they’re at the start/end.
/**
* Convert a string to title case, honoring common "small words".
*
* @param {string} str
* @param {Object} [opts]
* @param {string[]} [opts.smallWords]
* @returns {string}
*/
function toTitleCase(str, opts = {}) {
const defaultSmall = [
'a','an','the','and','but','or','nor','for','on','at','to','from','by','of','in'
];
const smallWords = new Set((opts.smallWords || defaultSmall).map(w => w.toLowerCase()));
return String(str)
.split(/\s+/)
.map((word, idx, arr) => {
const lower = word.toLowerCase();
const isEdge = idx === 0 || idx === arr.length - 1;
if (!isEdge && smallWords.has(lower)) return lower;
return lower.charAt(0).toUpperCase() + lower.slice(1);
})
.join(' ');
}
// Examples
// toTitleCase('a tale of two cities') -> 'A Tale of Two Cities'
// toTitleCase('from dusk till dawn') -> 'From Dusk Till Dawn'
Tip: This is a pragmatic approach—not a Chicago Manual of Style engine. For specific editorial rules, customize the smallWords list.
3) toCamelCase: Convert from delimiters to camelCase
Converts strings with spaces, dashes, or underscores into camelCase.
/**
* Convert a string to camelCase.
*
* @param {string} str
* @returns {string}
*/
function toCamelCase(str) {
const words = String(str)
.trim()
.replace(/['’]/g, '') // drop apostrophes in contractions
.replace(/[^A-Za-z0-9]+/g, ' ') // unify separators
.split(' ')
.filter(Boolean);
if (words.length === 0) return '';
const [first, ...rest] = words.map(w => w.toLowerCase());
return first + rest.map(w => w.charAt(0).toUpperCase() + w.slice(1)).join('');
}
// Examples
// toCamelCase('hello-world_test value') -> 'helloWorldTestValue'
// toCamelCase('API response code') -> 'apiResponseCode'
Pro tip: If you need to preserve acronyms (e.g., “API”), you can add a rule to keep uppercase if the word length <= 3 and all caps.
4) toKebabCase: Convert to kebab-case
Ideal for CSS classes, URLs, and CLI flags.
/**
* Convert a string to kebab-case.
*
* @param {string} str
* @returns {string}
*/
function toKebabCase(str) {
return String(str)
.trim()
.replace(/['’]/g, '')
.replace(/([a-z0-9])([A-Z])/g, '$1-$2') // handle camelCase input
.replace(/[^A-Za-z0-9]+/g, '-') // unify separators
.replace(/^-+|-+$/g, '') // trim dashes
.replace(/--+/g, '-') // collapse multiple
.toLowerCase();
}
// Examples
// toKebabCase('HelloWorldCSS') -> 'hello-world-css'
// toKebabCase(' Fancy Title!! ') -> 'fancy-title'
5) toSnakeCase: Convert to snake_case
Great for environment variables, database columns, or Python interop.
/**
* Convert a string to snake_case.
*
* @param {string} str
* @returns {string}
*/
function toSnakeCase(str) {
return String(str)
.trim()
.replace(/['’]/g, '')
.replace(/([a-z0-9])([A-Z])/g, '$1_$2')
.replace(/[^A-Za-z0-9]+/g, '_')
.replace(/^_+|_+$/g, '')
.replace(/__+/g, '_')
.toLowerCase();
}
// Examples
// toSnakeCase('ServerResponseTime') -> 'server_response_time'
// toSnakeCase('user name (primary)') -> 'user_name_primary'
6) slugify: URL-friendly slugs (with diacritics removal)
Generates clean slugs for SEO-friendly URLs. Removes diacritics, punctuation, and spaces.
/**
* Create a URL-friendly slug.
*
* @param {string} str
* @param {Object} [opts]
* @param {boolean} [opts.lowercase=true]
* @returns {string}
*/
function slugify(str, opts = {}) {
const { lowercase = true } = opts;
const s = String(str)
.normalize('NFD') // separate base letters and diacritics
.replace(/\p{M}/gu, '') // remove diacritic marks
.replace(/['’]/g, '') // drop apostrophes
.replace(/[^A-Za-z0-9]+/g, '-') // non-alnum -> dash
.replace(/^-+|-+$/g, '') // trim dashes
.replace(/--+/g, '-'); // collapse dashes
return lowercase ? s.toLowerCase() : s;
}
// Examples
// slugify('Crème Brûlée & Café!') -> 'creme-brulee-cafe'
// slugify('New Features in v2.0.0') -> 'new-features-in-v2-0-0'
Note: Uses Unicode property escapes; ensure your environment supports them (Node 12+ or modern browsers).
7) smartTruncate: Word-safe truncation with ellipsis
Truncates text to a limit, avoiding mid-word cuts and optionally preserving whole words.
/**
* Truncate a string to a length, optionally at word boundary, with suffix.
*
* @param {string} str
* @param {number} maxLength
* @param {Object} [opts]
* @param {string} [opts.suffix='…']
* @param {boolean} [opts.wordBoundary=true]
* @returns {string}
*/
function smartTruncate(str, maxLength, opts = {}) {
const { suffix = '…', wordBoundary = true } = opts;
const s = String(str);
if (s.length <= maxLength) return s;
const limit = Math.max(0, maxLength - suffix.length);
let cut = s.slice(0, limit);
if (wordBoundary) {
const lastSpace = cut.lastIndexOf(' ');
if (lastSpace > 0) cut = cut.slice(0, lastSpace);
}
return cut + suffix;
}
// Examples
// smartTruncate('This is a long sentence for preview text.', 20) -> 'This is a long…'
// smartTruncate('Supercalifragilisticexpialidocious', 10, { wordBoundary: false }) -> 'Supercali…'
Tip: For social previews or cards, aim for maxLength between 120–160 characters for readability.
8) stripTags: Remove HTML tags (for display-only use)
Quickly strip tags from HTML for display or indexing. Not a security sanitizer.
/**
* Strip HTML tags from a string. Not for sanitizing untrusted input.
*
* @param {string} html
* @returns {string}
*/
function stripTags(html) {
return String(html).replace(/<[^>]*>/g, '');
}
// Examples
// stripTags('<p>Hello <strong>world</strong>!</p>') -> 'Hello world!'
// stripTags('Line<br>Break') -> 'LineBreak'
Security note: Do NOT use regex tag stripping to prevent XSS. For sanitization, use a vetted library like DOMPurify (browser) or sanitize-html (Node).
Browser alternative (keeps text content):
function stripTagsSafe(html) {
const div = document.createElement('div');
div.innerHTML = html;
return div.textContent || '';
}
9) escapeHTML: Escape special HTML characters
Prevent HTML injection when embedding user input in HTML.
/**
* Escape &, <, >, " and ' for safe HTML rendering.
*
* @param {string} str
* @returns {string}
*/
function escapeHTML(str) {
return String(str)
.replace(/&/g, '&')
.replace(/</g, '<')
.replace(/>/g, '>')
.replace(/"/g, '"')
.replace(/'/g, ''');
}
// Examples
// escapeHTML(`<script>alert("x")</script> & me`) -> '<script>alert("x")</script> & me'
10) unescapeHTML: Decode escaped HTML entities
Reverse of escapeHTML for display or processing.
/**
* Unescape basic HTML entities.
*
* @param {string} str
* @returns {string}
*/
function unescapeHTML(str) {
return String(str)
.replace(/</g, '<')
.replace(/>/g, '>')
.replace(/"/g, '"')
.replace(/'/g, "'")
.replace(/&/g, '&'); // last to avoid double decoding
}
// Examples
// unescapeHTML('<h1>Title</h1>') -> '<h1>Title</h1>'
Note: This handles common named entities; for full HTML5 entity coverage, use a dedicated parser.
11) removeDiacritics: Strip accents/diacritics
Helpful for search indexing, comparisons, or generating ASCII-only IDs.
/**
* Remove diacritic marks from a string using Unicode normalization.
*
* @param {string} str
* @returns {string}
*/
function removeDiacritics(str) {
return String(str)
.normalize('NFD')
.replace(/\p{M}+/gu, '')
.normalize('NFC');
}
// Examples
// removeDiacritics('São Tomé e Príncipe') -> 'Sao Tome e Principe'
// removeDiacritics('Résumé') -> 'Resume'
12) isBlank: Detect empty or whitespace-only strings
A tiny guard for validation and early returns.
/**
* Check if a value is null/undefined or a string with only whitespace.
*
* @param {*} value
* @returns {boolean}
*/
function isBlank(value) {
if (value == null) return true;
return String(value).trim().length === 0;
}
// Examples
// isBlank(null) -> true
// isBlank(' ') -> true
// isBlank(' ok ') -> false
Tip: Use isBlank for simple validations and short-circuiting expensive work when inputs are empty.
13) countOccurrences: Count substring occurrences (with overlap)
Counts how many times a substring appears, with optional overlapping matches.
/**
* Count occurrences of a substring in a string.
*
* @param {string} str
* @param {string} search
* @param {Object} [opts]
* @param {boolean} [opts.overlap=false] - If true, allows overlapping matches.
* @param {boolean} [opts.caseSensitive=true]
* @returns {number}
*/
function countOccurrences(str, search, opts = {}) {
const { overlap = false, caseSensitive = true } = opts;
if (!search) return 0;
let s = String(str);
let q = String(search);
if (!caseSensitive) {
s = s.toLowerCase();
q = q.toLowerCase();
}
let count = 0;
let index = 0;
while (true) {
const found = s.indexOf(q, index);
if (found === -1) break;
count++;
index = found + (overlap ? 1 : q.length);
}
return count;
}
// Examples
// countOccurrences('banana', 'ana') -> 1
// countOccurrences('banana', 'ana', { overlap: true }) -> 2
// countOccurrences('Hello hello', 'hello', { caseSensitive: false }) -> 2
14) reverseGraphemes: Correctly reverse strings (emoji-safe)
Reverses by grapheme clusters, not just code units, so emoji and composed characters stay intact.
/**
* Reverse a string by grapheme clusters (emoji-safe).
*
* @param {string} str
* @param {string} [locale] Optional locale for segmenter
* @returns {string}
*/
function reverseGraphemes(str, locale) {
const s = String(str);
if (typeof Intl !== 'undefined' && Intl.Segmenter) {
const seg = new Intl.Segmenter(locale, { granularity: 'grapheme' }).segment(s);
const clusters = [];
for (const part of seg) clusters.push(part.segment);
return clusters.reverse().join('');
}
// Fallback: not perfect for all emojis (ZWJ sequences), but reasonable
return Array.from(s).reverse().join('');
}
// Examples
// reverseGraphemes('hello') -> 'olleh'
// reverseGraphemes('👍🏽💻') -> '💻👍🏽'
Why it matters: Naive reversal breaks sequences like family emojis or skin tone modifiers.
15) interpolate: Simple template replacement with {placeholders}
Lightweight templating for labels, messages, and logs. Supports nested paths.
/**
* Interpolate a string with {placeholders} using an object.
*
* @param {string} template
* @param {Record<string, any>} params
* @param {Object} [opts]
* @param {boolean} [opts.ignoreUnknown=true] - Leave tokens if not found
* @param {string} [opts.start='{']
* @param {string} [opts.end='}']
* @returns {string}
*/
function interpolate(template, params = {}, opts = {}) {
const { ignoreUnknown = true, start = '{', end = '}' } = opts;
function escapeRegExp(x) {
return x.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
}
function getByPath(obj, path) {
return path.split('.').reduce((acc, key) => (acc != null ? acc[key] : undefined), obj);
}
const re = new RegExp(`${escapeRegExp(start)}\\s*(.+?)\\s*${escapeRegExp(end)}`, 'g');
return String(template).replace(re, (_, token) => {
const val = getByPath(params, token);
if (val == null) return ignoreUnknown ? `${start}${token}${end}` : '';
return String(val);
});
}
// Examples
// interpolate('Hello, {user.name}!', { user: { name: 'Ava' } }) -> 'Hello, Ava!'
// interpolate('Total: {total} {currency}', { total: 19.99, currency: 'USD' }) -> 'Total: 19.99 USD'
// interpolate('Missing {value}', {}, { ignoreUnknown: false }) -> 'Missing '
Use cases:
- i18n strings with variables
- Log formatting
- Template-based filenames
Extra utilities you might also like
Depending on your project, these are handy too:
- padCenter: Center a string within a width using a fill character.
- dedent: Trim common leading indentation for multi-line strings.
- normalizeWhitespace: Collapse multiple spaces and newlines to a single space.
Here’s padCenter as a bonus:
/**
* Center-pad a string to a target length using a fill character.
*
* @param {string} str
* @param {number} targetLen
* @param {string} [fill=' ']
* @returns {string}
*/
function padCenter(str, targetLen, fill = ' ') {
const s = String(str);
if (s.length >= targetLen) return s;
const total = targetLen - s.length;
const left = Math.floor(total / 2);
const right = total - left;
return fill.repeat(left).slice(0, left) + s + fill.repeat(right).slice(0, right);
}
// Examples
// padCenter('Hi', 6, '-') -> '--Hi--'
Practical recipes: Combine utilities for real-world tasks
1) Build SEO-friendly blog post URLs
- Normalize the title, remove diacritics, and produce a clean slug.
const title = '10 Hidden Gems in São Paulo Nightlife';
const url = `/blog/${slugify(title)}`;
// -> '/blog/10-hidden-gems-in-sao-paulo-nightlife'
2) Safe preview text from rich content
- Remove HTML, trim, and truncate at word boundaries.
const html = '<p>Welcome to <strong>our platform</strong>. Enjoy your stay!</p>';
const preview = smartTruncate(stripTags(html), 50);
// -> 'Welcome to our platform. Enjoy your stay…'
3) Consistent identifiers from user input
- Convert display names to snake_case for system keys.
const display = 'Primary Contact (Billing)';
const key = toSnakeCase(display);
// -> 'primary_contact_billing'
4) Case conversions with diacritics handling
- Generate matching keys and URLs even with accented characters.
const label = 'Crème Brûlée';
const id = toCamelCase(removeDiacritics(label)); // 'cremeBrulee'
const link = slugify(label); // 'creme-brulee'
5) Message formatting with runtime variables
- Use interpolate for clean, maintainable strings.
const t = 'Hello {user.firstName}, you have {count} new {count,plural}!';
const msg = interpolate(t, { user: { firstName: 'Jo' }, count: 3 });
// -> 'Hello Jo, you have 3 new {count,plural}!' (unknown token left intact)
If you want smart pluralization, integrate a dedicated i18n/pluralization library; keep interpolate focused on key replacement.
Performance and correctness tips
- Prefer String.prototype methods and simple regex for hot paths; they’re fast and optimized in modern engines.
- Unicode matters:
- For accents: use normalize('NFD') + remove marks.
- For graphemes/emoji: use Intl.Segmenter when available.
- Avoid using regex for HTML sanitization. For untrusted input, always use a sanitizer.
- When counting or searching large texts, avoid constructing giant regex where simple indexOf loops suffice.
- Cache repeated conversions (like slugify for the same title) to reduce CPU work in list views.
Edge cases to consider
- Empty inputs: Each function above guards or returns sensible defaults.
- Locale-specific casing: Turkish dotted/dotless “i” is a classic pitfall; pass locale if you care.
- Over-aggressive normalization: Don’t remove diacritics unless you need ASCII-only or flexible search.
- Overlapping matches: Counting or replacing substrings can behave differently with overlap—be explicit.
- Emojis and combining characters: Reversal or capitalizing naive code units can break them; use grapheme-aware utilities where it matters.
Testing suggestions
- Add unit tests for:
- Mixed-case inputs: "HTMLParser", "user_name", "userName"
- Languages with diacritics: “Łódź”, “İstanbul”
- Emojis and ZWJ sequences: “👩💻”, “👨👩👧👦”
- Edge length constraints: boundary length equals suffix length
- Malformed HTML strings for stripTags
Example quick tests (Jest-like pseudo):
expect(slugify('Crème brûlée')).toBe('creme-brulee');
expect(toCamelCase('user_name-id')).toBe('userNameId');
expect(reverseGraphemes('👍🏽💻')).toBe('💻👍🏽');
expect(smartTruncate('word word', 6)).toBe('word…');
expect(countOccurrences('aaaa', 'aa', { overlap: true })).toBe(3);
Final thoughts
You don’t need a full utility library to handle everyday string tasks. With these 15 focused functions, you can:
- Normalize input consistently
- Generate clean URLs and identifiers
- Render user content safely
- Build friendly previews and titles
- Handle Unicode correctly when it counts
Copy what you need, adapt to your team’s conventions, and keep your string logic clear and predictable. When you outgrow these helpers, you’ll have a solid foundation to swap in specialized libraries without changing the shape of your code.