Skip to content

Parser helpers

Helpers for writing parsers (clijson.textutils). See Writing parsers for how they fit together.

textutils

Text helpers shared by all parsers.

These are deliberately small, well tested building blocks: cleaning captured output, converting values, slicing column-aligned tables and splitting output into blocks. Parser authors should reach for these before writing regexes.

clean_output

clean_output(text: str) -> str

Normalise newlines, drop ANSI escapes, pager prompts and backspaces.

dedent

dedent(text: str) -> str

Fast textwrap.dedent for right-stripped lines (spaces only).

blocks

blocks(
    text: str, start: str | Pattern[str] | None = None
) -> list[str]

Split text into blocks.

Without start, blocks are separated by blank lines. With start (a regex), a new block begins at each line matching it; text before the first match is discarded.

to_num

to_num(value: Any) -> Any

"42" -> 42, "1.5" -> 1.5, "1,024" -> 1024, anything else unchanged.

none_if

none_if(value: Any, *empty: str) -> Any

Return None for placeholder values like --, N/A or unassigned.

snake

snake(key: str) -> str

"Up Time (secs)" -> up_time_secs; "IP-Address" -> ip_address.

normalize_mac

normalize_mac(mac: str | None) -> str | None

Any MAC notation -> aa:bb:cc:dd:ee:ff; returns input if it is not a MAC.

parse_duration

parse_duration(value: str | None) -> int | None

Convert vendor uptime/age strings to seconds.

Handles 01:02:03, 1d02h, 3w4d, 2y10w, 5d 01:02:03, 1 week, 2 days, 3 hours, 4 minutes and plain second counts. Returns None for never/- or anything unrecognised.

header_columns

header_columns(
    header: str, names: Sequence[str] | None = None
) -> list[tuple[str, int]]

Return [(name, start_col), ...] for a header line.

Column names are split on 2+ spaces. When names is given, those exact strings are located in the header instead (useful for headers like Local Intf Holdtime where single spaces appear inside names).

slice_row

slice_row(line: str, starts: Sequence[int]) -> list[str]

Cut line at column starts, nudging cuts so words are never split.

parse_table

parse_table(
    text: str,
    header: str | Pattern[str] | None = None,
    names: Sequence[str] | None = None,
    keys: Sequence[str] | None = None,
    stop: str | Pattern[str] | None = None,
    skip: str | Pattern[str] | None = None,
    min_cells: int = 1,
    convert: bool = True,
    wrap: bool = False,
) -> list[dict[str, Any]]

Parse a column-aligned table.

Parameters:

Name Type Description Default
header str | Pattern[str] | None

regex locating the header line (default: first non-blank line)

None
names Sequence[str] | None

exact column titles as printed (default: split header on 2+ spaces)

None
keys Sequence[str] | None

output key names, defaults to snake(name)

None
stop str | Pattern[str] | None

regex; stop parsing at the first matching line

None
skip str | Pattern[str] | None

regex; ignore matching data lines

None
wrap bool

treat rows whose first cell is empty as continuations of the previous row

False

split_columns

split_columns(line: str, maxsplit: int = -1) -> list[str]

Split on runs of 2+ spaces (keeps single-spaced values together).

kv_pairs

kv_pairs(
    text: str,
    sep: str = "\\s*:\\s+|\\s*:\\s*$",
    key_re: str = "[A-Za-z][\\w \\-/().#'&]*?",
) -> dict[str, Any]

Extract Key: value pairs (several per line allowed, split on 2+ spaces or commas).

search

search(
    pattern: str | Pattern[str], text: str, flags: int = M
) -> dict[str, Any] | None

Regex search returning a dict of converted named groups (None values dropped).

match_lines

match_lines(
    pattern: str | Pattern[str], text: str, flags: int = 0
) -> Iterator[Match[str]]

Like re.finditer with re.M but a match can never span lines.

Prefer this over re.finditer(r"^...$", text, re.M): with \s+ between fields the latter silently glues a short line to the next one.

compact

compact(obj: Any) -> Any

Recursively drop None values and empty containers.