npm.io
0.17.0 • Published yesterday

yuku-parser

Licence
MIT
Version
0.17.0
Deps
2
Size
47 kB
Vulns
0
Weekly
0
Stars
942

yuku-parser

A fast, spec-compliant JavaScript and TypeScript parser, part of Yuku.

Install

npm install yuku-parser

It runs on Yuku's native core, installed for your platform. In browsers and edge runtimes, load the WebAssembly core from @yuku-core/wasm and pass it in as core.

Usage

import { parse } from "yuku-parser";

const { program, comments, diagnostics } = parse("const x = 1 + 2;");

The AST

ESTree for JavaScript and JSX, identical to Acorn, and TypeScript-ESTree for TypeScript, matching Oxc for both. On top of the specs, it carries stage 3 decorators, import defer and import source as phase on imports and ImportExpression, and the hashbang of Program.

Every node type is exported, listed in the type definitions.

import type { Expression, Identifier, Node, Statement } from "yuku-parser";

To walk, build, and check nodes, use yuku-ast. For scopes and bindings, use yuku-analyzer.

Options

parse(source, { path: "src/app.tsx" });
Option Values Default Description
path a file path none Names the file in diagnostics. lang and sourceType default from its extension.
lang "js", "jsx", "ts", "tsx", "dts" "js" The syntax to parse. .d.ts, .tsx, .ts, and .jsx paths select their own.
sourceType "module", "script", "commonjs" "module" "commonjs" allows top-level return, and .cjs and .cts paths select it.
preserveParens true, false true Keep ParenthesizedExpression nodes.
semanticErrors true, false false Also report the early errors that need scopes, such as redeclarations and break outside a loop.
attachComments true, false false Also attach each comment to its node. See Comments.
tokens true, false false Keep every token. See Tokens.
core a loaded core native The core that parses, see @yuku-core/wasm.

An unknown lang or sourceType throws a TypeError. langFromPath(path) and sourceTypeFromPath(path) return the values a path implies.

Result

interface ParseResult {
  program: Program;
  comments: Comment[];
  tokens?: TokenList; // with tokens: true
  diagnostics: Diagnostic[];
}

The parser recovers from errors, so a result with diagnostics still holds a tree of everything it could read.

Diagnostics

Every Yuku package reports diagnostics in one shape.

interface Diagnostic {
  severity: "error" | "warning" | "hint" | "info";
  message: string;
  path: string | null; // the path option
  start: number;       // UTF-16 offsets, like nodes
  end: number;
  labels: { start: number; end: number; message: string }[];
  help: string | null;
}

Comments

const { comments } = parse(`// a line comment\nconst x = 1; /* a block comment */`);
// [
//   { type: "Line", value: " a line comment", start: 0, end: 17 },
//   { type: "Block", value: " a block comment ", start: 31, end: 52 },
// ]

attachComments: true also hangs each comment on the node it sits next to, which yuku-codegen prints from, so comments move with their nodes.

const { program } = parse(`// header\nfunction foo() {} // trailing`, { attachComments: true });

program.body[0].comments;
// [
//   { type: "Line", position: "before", sameLine: false, value: " header" },
//   { type: "Line", position: "after", sameLine: true, value: " trailing" },
// ]

position is "before", "after", or "inside" an otherwise empty node, such as function f() { /* hi */ }.

Tokens

tokens: true keeps every token in a TokenList, a view over the parser's token table where a token is an index.

import { parse, TokenKind } from "yuku-parser";

const { tokens } = parse(source, { tokens: true });

for (let i = 0; i < tokens.length; i++) {
  if (tokens.kind(i) === TokenKind.Arrow) console.log(tokens.start(i), tokens.text(i));
}
tokens.kind(i)                      // one of TokenKind
tokens.text(i)                      // its source text
tokens.start(i)                     // UTF-16 offsets
tokens.end(i)

tokens.isKeyword(i)
tokens.isReserved(i)                // reserved unconditionally or in strict mode
tokens.isUnconditionallyReserved(i)
tokens.isStrictModeReserved(i)
tokens.isIdentifierLike(i)          // an identifier or any keyword
tokens.isNumericLiteral(i)
tokens.isBinaryOperator(i)
tokens.isLogicalOperator(i)
tokens.isUnaryOperator(i)
tokens.isAssignmentOperator(i)
tokens.precedence(i)                // binary precedence, 0 when none

tokens.newlineBefore(i)             // what ASI reads
tokens.escaped(i)
tokens.invalidEscape(i)             // a template chunk whose cooked value is undefined
tokens.loneSurrogate(i)

tokens.range(node)                  // [from, to) of the tokens inside a node
tokens.first(node)
tokens.last(node)
tokens.before(nodeOrOffset)         // the last token ending at or before it
tokens.after(nodeOrOffset)          // the first token starting at or after it
tokens.at(offset)                   // the token containing an offset

The queries answer with an index, or -1 when there is none. Tokens are as the parser resolved them, so a regex is one RegexLiteral, and comments are not tokens. The kinds are listed in tokens.d.ts.

License

MIT

Keywords