Rahul Juliato: Readable Regular Expressions for JavaScript/TypeScript, Inspired by Emacs' rx

Wait 5 sec.

IntroQuick, what does this match?/^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)(?:-((?:0|[1-9]\d*|\d*[a-zA-Z-][0-9a-zA-Z-]*)(?:\.(?:0|[1-9]\d*|\d*[a-zA-Z-][0-9a-zA-Z-]*))*))?(?:\+([0-9a-zA-Z-]+(?:\.[0-9a-zA-Z-]+)*))?$/;// Take// your// time...//// ...still decoding?//// OK, keep reading :)That's the official regexp from semver.org. Itvalidates version numbers like:// matches"1.2.3""0.10.0""2.0.0-rc.1""1.0.0-alpha.1+build.5""1.0.0+20260930"// doesn't match"01.2.3" // leading zero"1.2" // missing patch"v1.2.3" // no "v" prefix allowed"1.0.0-01" // numeric pre-release with a leading zero"1.2.3-" // empty pre-releaseDon't get me wrong, I love regexps, but in practice you probably spenda bunch of time writing one, testing it against some cases, and movingon, proud of your achievement!Some time passes and lucky future you (or unlucky someone else) has tochange it. Dramatic pause here.I bet you've been there. Now your options are probably: decode itagain from the start, rewrite the whole thing, or, in the age of AI,ask (and hopefully not blindly accept) an LLM for a new recipe.Emacs has had a nice answer for more readable regexps for a long time:the rx macro. I started using it all the time in Emacs Lisp, asreviewers always suggested it to me. Later, I started missing this DSLin JavaScript and TypeScript, so I wrote a small version of it for myprojects.So, what about reading that SemVer regexp like semver in the codebelow?const num = or("0", seq(anyOf("1-9"), zeroOrMore(digit)));const idChar = anyOf(alnum, "-");const preId = or(num, seq(zeroOrMore(digit), anyOf(alpha, "-"), zeroOrMore(idChar)));const dotted = (x: Item) => seq(x, zeroOrMore(".", x));const semver = RX( start, named("major", num), ".", named("minor", num), ".", named("patch", num), optional("-", named("pre", dotted(preId))), optional("+", named("build", dotted(oneOrMore(idChar)))), end,);The same strings match, and you get named groups as a bonus. By theend of this post you'll know every piece of it.TL;DR: jump straight to the cheat sheet,the side-by-side examples, the fullsource, or grab thegistto sneak a peek at the result.NOTE: the RX here has nothing to do withRxJS, which is an amazing library for reactiveprogramming with observables.A taste of rx in Emacs LispWith rx you describe a regexp as a tree of named forms, and Emacsturns it into the regexp string for you:(rx bos (+ digit) eos);; => "\\`[[:digit:]]+\\'"(rx bol "colo" (? "u") "r" eol);; => "^colou?r$"(rx bos "(" (= 3 digit) ")" space (= 3 digit) "-" (= 4 digit) eos);; => "\\`([[:digit:]]\\{3\\})[[:space:]][[:digit:]]\\{3\\}-[[:digit:]]\\{4\\}\\'"A few things to notice:Strings are literals. "(" means a parenthesis. You don'tneed to escape anything by hand.Sequence is implicit. Every form takes a list of things andmatches them one after the other. You don't need to wrap them in aseq, even though seq exists.Groups appear only when needed. (+ digit) becomes[[:digit:]]+, not \(?:[[:digit:]]\)+.The proposed JavaScript/TypeScript version in this post reads likethis:const phone = RX( start, "(", repeat(3, digit), ")", space, repeat(3, digit), "-", repeat(4, digit), end,);// => /^\(\d{3}\)\s\d{3}-\d{4}$/Under the hoodIf you want strings to be literals, you can't represent a regexp pieceas a plain string, otherwise you can't tell "(" (a literalparenthesis) apart from "(?:...)" (a group you built). So everypiece is a small object:type Kind = "atom" | "seq" | "alt";interface RxNode { readonly src: string; readonly kind: Kind; readonly set?: string; // char sets only, see below readonly neg?: boolean;}type Item = string | RxNode;src is the regexp text. kind records how that text behaves whenyou glue it to other things:atom: a single unit, like a, \d, [a-z] or (...). You canput a quantifier right after it.seq: safe to concatenate, but a quantifier needs (?:...) aroundit. abc is a seq, and so is a+, since a+? would silentlyturn into a lazy quantifier.alt: has a | at the top level, so it needs (?:...) almosteverywhere.Plain strings go through literal, which escapes them:const esc = (s: string) => s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");const literal = (s: string): RxNode => ({ src: esc(s), kind: s.length === 1 ? "atom" : "seq",});const toNode = (x: Item): RxNode => (typeof x === "string" ? literal(x) : x);With that in place, seq joins nodes and only brackets alternations:const seq = (...xs: Item[]): RxNode => { const nodes = xs.map(toNode).filter((n) => n.src !== ""); if (nodes.length === 0) return { src: "", kind: "seq" }; if (nodes.length === 1) return nodes[0]; let src = ""; for (const n of nodes) {const part = n.kind === "alt" ? `(?:${n.src})` : n.src;// `\1` followed by a literal `0` would read as `\10`if (/\\\d+$/.test(src) && /^\d/.test(part)) src += "(?:)";src += part; } return { src, kind: "seq" };};(That backreference check is one of those bugs you only find bywriting tests, or when it happens to you in prod. backref(1)followed by the literal "0" gives you backreference number ten.)Every quantifier is a seq of its arguments plus a suffix, bracketedonly when the body isn't an atom:const quantifiable = (n: RxNode) => n.kind === "atom" ? n.src : `(?:${n.src})`;const quantifier = (suffix: string) => (...xs: Item[]): RxNode => ({src: quantifiable(seq(...xs)) + suffix,kind: "seq", });const zeroOrMore = quantifier("*");const oneOrMore = quantifier("+");const optional = quantifier("?");Because each quantifier calls seq on its arguments, you get theimplicit sequence for free: optional("-", group(x)) becomes(?:-(x))?.And finally, the two entry points. As in Emacs, rx returns astring. RX returns a RegExp you can use right away:const rx = (...xs: Item[]): string => seq(...xs).src;function RX(...xs: Item[]): RegExp { return new RegExp(rx(...xs));}RX.flags = (flags: string, ...xs: Item[]): RegExp => new RegExp(rx(...xs), flags);RX.flags exists because Emacs controls case folding through thecase-fold-search variable, and JavaScript puts it on the regexpitself.That's the whole engine! Now, let's build our vocabulary.Character setsIn Emacs you write (any "a-z" "_"). Inside those strings, a-z is arange, and a - at either end is a plain dash. I kept the same rule:const hexDigit = anyOf("0-9a-fA-F");RX(start, "#", repeat(6, hexDigit), end);// => /^#[0-9a-fA-F]{6}$/RX(start, optional(anyOf("+-")), oneOrMore(digit), end);// => /^[+\-]?\d+$/The dash comes out escaped because sets can merge. If you combineanyOf("+-") with anyOf("0-9"), an unescaped - would end up inthe middle and create a range from + to 0. Escaping it costs onebackslash.And merging is the reason why RxNode has a set field. It is thereto hold the text that goes between [ and ], so anyOf can takeother sets as arguments:const lower = anyOf("a-z");const upper = anyOf("A-Z");const alpha = anyOf(lower, upper); // [a-zA-Z]const alnum = anyOf(alpha, "0-9"); // [a-zA-Z0-9]not negates a set, and it knows the shorthand classes:not(digit); // \Dnot(anyOf(space, "@")); // [^\s@]notChar(","); // [^,] (rx's not-char)The simple email check, which most of us have written as/^[^\s@]+@[^\s@]+\.[^\s@]+$/ at some point, becomes:const part = oneOrMore(not(anyOf(space, "@")));RX(start, part, "@", part, ".", part, end);// => /^[^\s@]+@[^\s@]+\.[^\s@]+$/The rest of the Emacs character classes are there too: digit,hexDigit, space, blank, wordChar, notWordChar, alpha,alnum, lower, upper, punct, control, graphic, printing,ascii and nonascii. One difference: in Emacs they understandUnicode, and mine are ASCII only. alpha won't match é.Two more come from rx's symbol list, and people (me, many times) mixthem up:const notNewline: RxNode = { src: ".", kind: "atom" }; // rx: nonlconst anything = set("\\s\\S"); // rx: anything / anycharIn rx, anything really means anything, newlines included. Here iswhere the difference shows up:const code = "a = 1; /* first\n second */ b = 2;";RX("/*", zeroOrMoreLazy(notNewline), "*/").exec(code);// => nullRX("/*", zeroOrMoreLazy(anything), "*/").exec(code)?.[0];// => "/* first\n second */"Alternatives, and the longest matchor works as you'd expect, and gets bracketed when it lands inside asequence:RX(start, or("cat", "dog", "bird"), end);// => /^(?:bird|cat|dog)$/Did you notice the order changed? I copied that behavior fromEmacs. When every branch of an or is a plain string, rx hands themto regexp-opt, which builds a pattern that prefers the longestmatch:(rx (or "in" "int" "interface"));; => "\\(?:in\\(?:t\\(?:erface\\)?\\)?\\)"JavaScript alternation takes the first branch that matches, going leftto right. So the naive regexp for a list of keywords has a 'bug':/in|int|interface/.exec("interface Foo")?.[0];// => "in"RX(or("in", "int", "interface")).exec("interface Foo")?.[0];// => "interface"I don't build a trie like regexp-opt does. Sorting the strings bylength, longest first, is enough to get the same behavior:const or = (...xs: Item[]): RxNode => { if (xs.length === 0) return unmatchable; if (xs.length === 1) return toNode(xs[0]); const branches = xs.every((x) => typeof x === "string")? [...(xs as string[])].sort((a, b) => b.length - a.length): xs; return { src: branches.map((x) => toNode(x).src).join("|"), kind: "alt" };};As in Emacs, or() with no branches returns unmatchable, which is(?!) here. It's handy when you build the branch list at runtime andit might come out empty.Repetition, greedy and lazyEmacs has (= n ...), (>= n ...) and (** n m ...). Here they arerepeat, atLeast and between:RX(start, between(2, 4, digit), end); // /^\d{2,4}$/RX(atLeast(3, digit)); // /\d{3,}/The lazy versions *?, +? and ?? are zeroOrMoreLazy,oneOrMoreLazy and optionalLazy. The classic HTML tag example:const html = "bold and italic";RX("").exec(html)?.[0];// => "bold and italic"RX("").exec(html)?.[0];// => ""Groups and backreferencesgroup is a capturing group, and backref points back to it:RX(start, group(oneOrMore(wordChar)), space, backref(1), end);// => /^(\w+)\s\1$/ matches "hello hello", not "hello world"Emacs also has (group-n N ...) to pick the group number. JavaScriptcan't do that, but it has named groups, which serve the same purposeand read better:const date = RX( named("y", repeat(4, digit)), "-", named("m", repeat(2, digit)), "-", named("d", repeat(2, digit)),);date.exec("2026-09-30")?.groups;// => { y: '2026', m: '09', d: '30' }backref accepts a name as well:RX( "", zeroOrMoreLazy(notNewline), "",);// => /.*?/Anchorsrx distinguishes the start of the string (bos) from the start of aline (bol). In JavaScript both are ^, and the m flag decideswhich one you get. I kept both names so the intent shows in the code:const text = "TODO: write post\nDONE: fix rx\nTODO: publish";const todo = RX.flags( "gm", lineStart, "TODO: ", named("task", oneOrMore(notNewline)), lineEnd,);[...text.matchAll(todo)].map((m) => m.groups?.task);// => [ 'write post', 'publish' ]Why not add start and end automatically? Because you only wantthem when validating a whole string. When searching inside a text, asin split, replace or matchAll, a hidden ^ and $ would breakeverything. Emacs agrees: bos and eos are explicit in rx too.wordBoundary and notWordBoundary map straight to \b and \B.Emacs also has bow and eow (\), start and end of aword. JavaScript lacks those, so I combined \b with a lookaround:const wordStart = zeroWidth("\\b(?=\\w)");const wordEnd = zeroWidth("\\b(?", zeroOrMoreLazy(notNewline), "",);// "bold" -> true// "oops" -> falseWhole wordcat as a word, not inside another one.// Emacs: (rx word-boundary "cat" word-boundary)const regex = /\bcat\b/;const dsl = RX(wordBoundary, "cat", wordBoundary);// "the cat sat" -> true// "concatenate" -> falseCase-insensitivehello, in any case.// Emacs: (let ((case-fold-search t))// (string-match-p (rx bos "hello" eos) "HeLLo"))const regex = /^hello$/i;const dsl = RX.flags("i", start, "hello", end);// "HeLLo" -> true// "help" -> falseFull sourceIt's a single file with no dependencies. Copy it into your projectand start deleting the forms you don't need, or adding the ones youmiss.You can check the same code, plus all the examples from this post (anda few more), in thisgist.If you'd rather not set anything up, paste it into the TypeScriptPlayground, hit "Run", and checkthe "Logs" tab./* ========================================================= * CORE * ========================================================= */// How a node behaves when combined with others:// atom -> single unit, a quantifier can be glued right after it// seq -> safe to concatenate, needs (?:) to be quantified// alt -> has a top-level `|`, needs (?:) almost everywheretype Kind = "atom" | "seq" | "alt";interface RxNode { readonly src: string; readonly kind: Kind; // char sets only: the text that goes inside [ ], so sets can merge readonly set?: string; readonly neg?: boolean;}type Item = string | RxNode;const esc = (s: string) => s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");const escSet = (s: string) => s.replace(/[\]\\^-]/g, "\\$&");// rx: (literal EXPR) — a string computed at runtime, matched as-isconst literal = (s: string): RxNode => ({ src: esc(s), kind: s.length === 1 ? "atom" : "seq",});const toNode = (x: Item): RxNode => (typeof x === "string" ? literal(x) : x);// wrap a node so a quantifier applies to all of itconst quantifiable = (n: RxNode) => n.kind === "atom" ? n.src : `(?:${n.src})`;// rx: (regexp EXPR) — escape hatch, trust the regexp as-isconst raw = (re: string | RegExp): RxNode => ({ src: typeof re === "string" ? re : re.source, kind: "alt",});/* ========================================================= * COMPOSITION * ========================================================= */const seq = (...xs: Item[]): RxNode => { const nodes = xs.map(toNode).filter((n) => n.src !== ""); if (nodes.length === 0) return { src: "", kind: "seq" }; if (nodes.length === 1) return nodes[0]; let src = ""; for (const n of nodes) {const part = n.kind === "alt" ? `(?:${n.src})` : n.src;// `\1` followed by a literal `0` would read as `\10`if (/\\\d+$/.test(src) && /^\d/.test(part)) src += "(?:)";src += part; } return { src, kind: "seq" };};const unmatchable: RxNode = { src: "(?!)", kind: "atom" };// Like rx: when every branch is a plain string, try the longest first,// so or("in", "int") matches "int" instead of stopping at "in".const or = (...xs: Item[]): RxNode => { if (xs.length === 0) return unmatchable; if (xs.length === 1) return toNode(xs[0]); const branches = xs.every((x) => typeof x === "string")? [...(xs as string[])].sort((a, b) => b.length - a.length): xs; return { src: branches.map((x) => toNode(x).src).join("|"), kind: "alt" };};/* ========================================================= * CHARACTER SETS * ========================================================= */const set = (body: string, neg = false): RxNode => ({ src: neg ? `[^${body}]` : `[${body}]`, kind: "atom", set: body, neg,});// class escapes are sets too, so they can go inside anyOf(...)const classEscape = (e: string): RxNode => ({ src: e, kind: "atom", set: e, neg: false,});// Same reading as rx: inside a string, "a-z" is a range, while a `-`// at the start or end is just a dash ("+-" is plus or minus).const intervals = (s: string): string => { let body = ""; let i = 0; while (i { if (typeof x === "string") return intervals(x); if (x.set === undefined || x.neg)throw new Error(`anyOf: not a positive char set: ${x.src}`); return x.set;}).join(""); return set(body);};// rx: (not charset) — not(digit) -> \D, not(anyOf(",;")) -> [^,;]const not = (x: Item): RxNode => { const n = typeof x === "string" ? anyOf(x) : x; if (n.set === undefined) throw new Error(`not: not a char set: ${n.src}`); if (n.neg) return set(n.set); if (/^\\[dswDSW]$/.test(n.src)) {const c = n.src[1];const flipped = c === c.toLowerCase() ? c.toUpperCase() : c.toLowerCase();return classEscape(`\\${flipped}`); } return set(n.set, true);};// rx: (not-char "a-z" ...) — shorthand for (not (any ...))const notChar = (...xs: Item[]) => not(anyOf(...xs));// rx char classes, `[[:name:]]` in Emacsconst digit = classEscape("\\d");const space = classEscape("\\s");const wordChar = classEscape("\\w");const notWordChar = not(wordChar);const lower = anyOf("a-z");const upper = anyOf("A-Z");const alpha = anyOf(lower, upper);const alnum = anyOf(alpha, "0-9");const hexDigit = anyOf("0-9a-fA-F");const blank = set(" \\t");const control = set("\\x00-\\x1f\\x7f");const punct = anyOf("!-/:-@[-`{-~");const graphic = anyOf("!-~");const printing = anyOf(" -~");const ascii = set("\\x00-\\x7f");const nonascii = set("\\u0080-\\uffff");// rx: `nonl` is any char but newline; `anything` really is anythingconst notNewline: RxNode = { src: ".", kind: "atom" };const anything = set("\\s\\S");/* ========================================================= * ANCHORS (zero-width) * ========================================================= */const zeroWidth = (src: string): RxNode => ({ src, kind: "seq" });// rx: bos / eosconst start = zeroWidth("^");const end = zeroWidth("$");// rx: bol / eol — same symbols, only per line with the "m" flagconst lineStart = start;const lineEnd = end;const wordBoundary = zeroWidth("\\b");const notWordBoundary = zeroWidth("\\B");// rx: bow / eow — JS has no \< \>, so a boundary plus a lookaroundconst wordStart = zeroWidth("\\b(?=\\w)");const wordEnd = zeroWidth("\\b(?