Advanced usage
Deeper mechanics behind features the standalone viewer and Python/Jupyter library guides only introduce — for when you need to know why something behaves the way it does, not just how to call it. See Examples & study material for runnable code; this page is about the internals those examples exercise.
Table of contents
- Table of contents
- Diff views and table alignment
- Starting a diff
- Why Split may not use each file's literal table order
- Embedded mode (the host↔viewer protocol)
- Renderers: how they actually work
- Interpreter selection
- Interpreter pairing in the standalone viewer
- Distributing interpreters as a package
convert()'s built-in/JS-runtime boundary- Config resolution, in full
- Command line reference
Diff views and table alignment
Diff mode offers two presentations of the same comparison and three scopes. The info badge beside Diff in the viewer header provides the same guide without leaving the page.
Starting a diff
The Diff… button in the header is hidden until real data is loaded
(hasLoadedRealData/updateDiffButtonState) — comparing the empty default
placeholder against something else isn't a meaningful diff. Three ways to
reach it, all funneling into the same file-intake logic:
- Click Diff…, then pick file(s) for the other side — the already-open document becomes the fixed left side (re-used as-is, not re-parsed).
- Drag two files onto the page at once (
dispatchFiles). Dropping one plain data file, or one markup file together with its interpreter, still just loads it normally — not a diff. Two data files (in any combination of formats, plus at most one shared interpreter) start a diff intake instead; which file becomesLeft:vsRight:follows drop/selection order, confirmed on-screen by those exact labels in the diff header. Formats don't have to match: a.jsonand an.xml+ interpreter start a cross-format diff too, entirely client-side, no Python involved —dispatchFilesresolves each side's format independently and never compares them. (The same cross-format diff is also reachable from Python — seestructile.diff().) Anything else (three or more files, files that don't classify as data or interpreter) is rejected with a note explaining what's expected, rather than guessing. - Drop file(s) straight onto the Diff… button itself — same effect as clicking it first, then dropping.
Assembling both sides doesn't have to happen in one drop. The
diffIntake state machine accumulates a data file per side plus a pool of
interpreter candidates across as many separate drops/picks as it takes, in
any order — the #diffIntakePanel status box (renderDiffIntakeStatus)
shows what's still missing per side (e.g. "Right: needs an interpreter
(.js) — tried 1, none worked yet"), and every new interpreter drop is
retried against every unresolved side. Dropping a second data file onto a
side that's already filled is rejected with a specific message instead of
silently replacing it; Choose file(s)…/Cancel in the panel work
alongside drag-and-drop for the same intake.
Once a diff is open:
- ⇄ (
swapDiffSides) swaps left and right — inverts added/removed and old/new, keeps matched rows/keys stable. - Click a
Left:/Right:name in the diff header (startDiffReplaceSide) to replace just that side: a fresh intake opens with the other side fixed, seeded with both sides' current interpreters as candidates (the replacement might turn out to match either schema).
Try it with no Python at all — every samples/diff_demo*.html file is a
pre-built, self-contained standalone diff page, directly double-click-
openable: diff_demo.html (Python-repr,
including Set diffing), diff_demo_chains.html,
diff_demo_config.html,
diff_demo_tables.html (same-format
JSON pairs), and the two cross-format examples also linked from the
README's Python diff walkthrough,
diff_demo_tables_xml.html (XML vs.
JSON) and diff_demo_tagged_ab.html
(two mutually-incompatible XML schemas, one interpreter per side) —
diff_demo_tagged_ac.html and
diff_demo_tagged_bc.html round
that out with a third, equally incompatible schema
(interpreters/keyed_xml.js — see
"Interpreter selection" below), each diffed against both of the others.
| Scope | Split | Unified |
|---|---|---|
| Full | Complete left and right data in aligned panes. One-sided data is red on the left and green on the right; the missing-side copy is struck through. Values present on both sides but changed are amber in both panes. | One complete superset. Added data is green, removed data red, and an old → new value is amber at one logical location. |
| Compact | The innermost changed containers in two panes, with breadcrumbs and complete unchanged local context. | The same changed containers and context in one superset. |
| Ultracompact | The same containers and breadcrumbs as Compact, pruned to changes plus the identifiers and order markers needed to interpret them. | The same pruning in one superset. |
Why Split may not use each file's literal table order
Split deliberately normalizes matching tables so both panes have identical row and column positions. Matched rows and columns follow the left file's order. Rows or columns that exist only on the right are appended afterward, in their relative order from the right file. This makes horizontal visual comparison reliable instead of placing the same logical cell at different coordinates.
For example, samples/sample_tables.json lists OIL before USD in the
correlations matrix, while samples/sample_tables_right.json lists USD
before OIL. Split renders both panes as OIL, then USD, because the left
file defines the matched order. The comparison still records and displays the
order change through order markers: a distinct accent-colored outline
(.diff-order — an inset border, not a red/green/amber fill) drawn around a
reordered table row/column header or a reordered dict/table's own name,
separate from the add/remove/changed coloring above — alignment does not
erase that difference, it just relocates it from "position" to "marker."
Plain dictionary keys and set members use their canonical diff order rather than table row/column alignment. Search is available in both presentations; in Split it searches and navigates matches in both panes.
Embedded mode (the host↔viewer protocol)
A host page can embed structile.html in an <iframe> (or a
srcdoc document, as the VS Code extension does) as a controlled editing
surface — the host supplies the data/config instead of the user picking a
file, and the viewer reports state back instead of downloading
files itself. This is what structile's widget renderer and the VS
Code extension both build on. See the EMBEDDED MODE section near the top
of the <script> in structile.html for the authoritative
implementation and comments; this is a summary.
Turning it on — either works, and either activates it independent of the other:
- Load the page with
?embed=1in the URL, or - Just post a
structile-embed-initmessage to it (useful when you can't control the iframe'ssrcquery string, e.g. when loading it viasrcdoc).
What changes: file load, paste, view save/load, and the about popover
are hidden — everything the host now owns. The canvas, undo/redo, and
search stay. The settings panel is hidden the same way UNLESS the host
opts in via enableSettingsPanel (see below) — both structile's widget
and the VS Code extension do.
Editing now matches the standalone viewer exactly for value/table-cell
edits: double-clicking one opens the same reviewable Edit diff (original on
the left, your edit on the right — see enterSelfEditMode in
structile.html) instead of committing immediately with a small
red dirty-dot indicator. Save (the diff header's own Save button, or a
host's pull-save via structile-embed-request-save — see hostManagesSave below)
commits it back; Cancel discards it. Key/column/name renames are the one
deliberate exception — they still commit immediately, live, with the old
dirty-dot indicator, since the Edit-review diff has no graph-based rename
affordance (a renamed key would have to resolve against a synthesized,
compound "old → new" display key — genuinely ambiguous to get right, the
same reason standalone's own Edit mode requires typing a rename into its
source pane instead of double-clicking one; a disableSourceView host has
no source pane to type into at all). structile-dirty-state-changed still fires
for both kinds of edit — a value edit under review pushes it from the diff
engine's own added/removed/changed counts (not per-field dirty dots, which
don't apply once the graph is a diff-engine tree), a rename still pushes it
the original way.
Message protocol (plain objects, matched on type):
| Direction | Type | Payload |
|---|---|---|
| host to viewer | structile-embed-init |
{ value? \| text?+format?, name?, interpreterSource?, interpreterCandidateNames?, config?, hostManagesSave?, disableSourceView?, enableSettingsPanel? } |
| host to viewer | structile-embed-diff-init |
{ left:{text,format,name?}, right:{text,format,name?}, leftInterpreterSource?, rightInterpreterSource?, config?, keyColumns?, diff?:{view}, hostManagesSave?, disableSourceView? } — two-sided diff mode (structile.diff(renderer="widget")); see below, and Diff views and table alignment |
| host to viewer | structile-embed-interpreter |
{ source, name? } — send a companion interpreter after the fact |
| host to viewer | structile-embed-request-save |
{ forBackup?: boolean } — pull current content on the host's own initiative (see hostManagesSave below); forBackup:true (VS Code's periodic hot-exit snapshot, not a real save) skips the usual dirty-clearing/edit-review-commit side effects a normal pull-save has |
| viewer to host | structile-dirty-state-changed |
{ dirty, totalCount, containers: [{path, label, count}] } — single-value mode only, never sent for diff mode |
| viewer to host | structile-save-requested |
{ format, ext, fileName, isSaveAs:false, content } | { format, ext, fileName, isSaveAs:false, content, side } (diff mode — side is "left"/"right") | { error } |
| viewer to host | structile-settings-changed |
{ settings, renderMap, links, theme } — only from the settings panel's "Set as default" button (requires enableSettingsPanel); see below |
| viewer to host | structile-graph-save-requested |
no payload — a hostManagesSave host hides the viewer's own Save/Save As buttons entirely, so this just asks the host to run its own save command (e.g. Ctrl+S reaching VS Code's native save) instead of the viewer pulling/serializing anything itself |
| viewer to host | structile-embed-interpreter-resolved |
{ candidates: [{name, outcome, error?}] } — reported whenever an interpreter trial actually runs (single candidate or a list); outcome is "used" (the one that won), "raised" (threw — error has its message), or "not_tried" (never reached because an earlier one already won). name comes from the matching interpreterCandidateNames entry sent in on structile-embed-init, or "candidate N" (1-based) when that wasn't supplied. Deliberately host-agnostic — the VS Code extension's Structile: Show Interpreter Resolution command/status bar are its first consumer, but any embedder can use it |
For structile-embed-init: pass either an already-parsed value, or raw text
plus format — "json" and "python" use the two built-in parsers;
anything else ("xml", "html", or a custom format like "ini") runs
through an interpreter (DOM-based for "xml"/"html", raw-text for
anything else — see Interpreter selection).
config is a settings payload in the same shape as a .settings.json file
(e.g. { settings: { valMax: 22, n: 3 } }), plus an optional top-level
theme: "light" | "dark" to set the initial theme, and an optional
links: { open?, close? } — a dict value/table cell wrapped between
open/close, matched verbatim (delimiters included), renders as a
clickable link that jumps to whichever dict/table elsewhere in the
document has that exact same name (off unless open is set; see
linkOpen/linkClose in structile/config.py). text sent with a
non-json/python format but without interpreterSource yet just waits for
a later structile-embed-interpreter message, the same pairing UX a
drag-and-dropped file uses (see "Interpreter pairing" below) — regardless
of whether the format is markup or not. interpreterCandidateNames: string[]
optionally labels each interpreterSource candidate (only meaningful when
interpreterSource is an array, or to give a lone candidate a real name
instead of a generic placeholder) — purely for reporting which one won or
lost a trial back to the host via structile-embed-interpreter-resolved
below; it never affects which candidate is tried or in what order. Omit it
and reported names fall back to "candidate N".
hostManagesSave: true hides the Save button and stops
Ctrl/Cmd+S from being handled inside the iframe at all — for a host (like
the VS Code extension) that wants its own save keybinding to win instead;
that host is expected to pull current content on demand via
structile-embed-request-save rather than waiting for a push.
disableSourceView: true hides the read-only raw-source-text toggle/pane
entirely — the VS Code extension's own
custom editor sets this (it gets its own, separately-designed source view
later); the structile widget's embed-init never sets it, so the toggle
is present there by default, same as standalone-page use.
enableSettingsPanel: true shows the settings panel/button instead of the
default fully-hidden embed behavior — both structile's widget and the
VS Code extension set this. Its embed-only "Set as default" button (next to
Export/Load/Reset settings) posts structile-settings-changed with the panel's
FULL current state — buildSettingsFile()'s own {settings, renderMap,
links} output plus theme — rather than applying only to this one view; each host
persists it as its own new baseline for views it creates from then on:
structile.options (in-process, matplotlib.rcParams-style — see
replace_options in structile/config.py) for the widget,
structile.* Workspace (falling back to User) settings for the VS
Code extension (applySettingsAsDefault in vscode-extension/src/format.ts).
It's a one-off action, not continuously-observed state, so — like
structile-save-requested and unlike structile-dirty-state-changed — a host should
treat it as an event to react to once, not a value to read back later.
structile-embed-diff-init (used by structile.diff(renderer="widget") —
StructileDiffWidget in widget.py) puts the viewer into the same
two-sided diff state a standalone window.__STRUCTILE_SNAPSHOT__ with
mode:"diff" does (see Diff views and table alignment
above). The diff GRAPH stays
permanently read-only in both cases (no graph editing, undo/redo, or a
second diff) — but each side's own source pane is independently
editable/saveable whenever the source view itself isn't disabled, via its
own Save button (performDiffSideSave, posting structile-save-requested with
side). Both sides must share one format in this version; a markup
(xml/html) side uses leftInterpreterSource/rightInterpreterSource
(a single source or a list of candidates, same as structile-embed-init's own
interpreterSource) rather than a single shared one, since the two sides
can be different, mutually-incompatible schemas.
In embedded mode, Save never touches the filesystem or opens a tab — it
only posts structile-save-requested with the serialized content, and it's the
host's job to persist it. The compatibility isSaveAs field is always
false; VS Code can still offer its own native Save As command because it
chooses the destination before pulling content from the viewer.
Note: the viewer posts with
targetOrigin: "*"(it doesn't know the host's origin ahead of time) and listens for messages from any origin — fine for a same-app iframe/widget, but don't embed untrusted third-party pages this way without adding your own origin checks on both ends.
Renderers: how they actually work
structile.render (see Examples & study material —
02_renderers.ipynb in the study zip — for the user-facing side of this)
never has separate code paths per
renderer for building content — every renderer is handed the same
Payload (text, format, name, interpreter_source?, config?) and
either:
- drives the embedded-mode protocol above (
widget), or - builds ONE self-contained standalone HTML page from it (
browser/file, andRenderHandle.to_html()/.open()/.save()for every renderer, includingwidget— the same page can always be popped out later).
Building the standalone page: render.build_standalone_html() reads
structile.html verbatim, then injects one small
<script>window.__STRUCTILE_SNAPSHOT__ = {...};</script> block immediately
before the file's own main <script> tag (found via an exact-string
anchor, not a fragile regex — see the VS Code extension guide for the "CSP
script-tag targeting fragility" lesson behind this, which is exactly why
this isn't a naive substring search either). structile.html's own
init() checks for window.__STRUCTILE_SNAPSHOT__ and, if present, feeds it
into applyInitialPayload() — the same function structile-embed-init uses —
but without calling enterEmbedMode(), so none of the standalone
chrome (settings bar, Save) gets hidden. That's the
entire difference between a browser/file render and an embedded one:
which function gets called around the same payload-loading logic.
Saving a standalone page always asks the browser user for a destination,
writes/downloads the serialized data, and opens that saved content in a new
clean viewer tab. The clean tab is built from the viewer document with the
window.__STRUCTILE_SNAPSHOT__ bootstrap removed, then populated through a
nonce-scoped BroadcastChannel handoff. The handoff is checked before any
snapshot during initialization, so a sample HTML page cannot reload its
integrated original over the just-saved content.
Why the payload is always text+format, never a pre-parsed value:
structile.html preserves integers outside JavaScript's safe
range (|v| >= 2**53) by re-scanning the text of a JSON document for
bare integer literals as it parses (jsonParsePreserveBigInt) — a
JS-object-literal embedding would just silently round through a double
the moment V8 parsed the surrounding <script> block. Routing everything
through text+format="json" (i.e. jsonParsePreserveBigInt(text), not a
directly-embedded object) is what keeps huge integers exact all the way
through a browser/file render too, not just widget.
The session temp directory: browser/file write into
tempfile.mkdtemp(prefix="structile_"), created lazily on first use and
reused for the rest of the process — and never cleaned up. That's
deliberate: the browser may not have finished loading the file by the time
the Python process exits, and often you want to keep the snapshot around
afterward anyway.
Interpreter selection
See Choosing obj, format, and interpreter for a practical,
case-by-case guide to how obj/format=/interpreter= combine when
calling open()/diff(). This section is the algorithm underneath
that — how a candidate list is built and picked from, not how to invoke
it.
structile.interpreters is generic across every markup extension on
purpose — the only place "XML" and "HTML" mean anything different at all is
MARKUP_FORMATS (which DOMParser mode an extension implies); the
registry, candidate-list resolution, and try-in-order selection below never
branch on which one they're looking at.
An interpreter format is either markup ("xml"/"html", an entry in
MARKUP_FORMATS) or anything else ("ini", or any other string) — the
former runs the interpreter's interpretXML(xmlDocument)/serializeXML(value)
against a browser-parsed DOM document, the latter runs
interpretText(text)/serializeText(value) against the raw text directly,
no DOM involved. Which contract applies is decided purely by the format
label, in exactly one place client-side
(tryInterpreterCandidates/structuredFormat in structile.html)
— everything below (the registry, resolve_candidates,
select_interpreter) is agnostic to which contract a given format actually
uses.
The registry (register_interpreter/unregister_interpreter/
get_registered_interpreter) is a plain dict[extension, spec], where
spec is a single path/raw-source or a list of them. resolve_candidates(explicit, ext)
normalizes whichever one applies (an explicit interpreter= always wins
over the registry) into a plain list — this is pure and does no I/O, so
"an interpreter registered for .html is never touched while resolving
.xml" is true by construction, not by convention (verified by this project's own test suite,
not just by convention).
Selecting among candidates (select_interpreter) tries each in order
and returns the first one that doesn't raise — but open() itself
never calls it at all anymore. Usage must never require running any
JS on the Python side (standalone HTML, Jupyter, the Python API, the VS
Code extension — some of these run on machines where installing anything
beyond structile's own dependencies isn't even an option), so every
open() renderer — widget included — forwards interpreter
source(s) as-is and lets the browser try them:
- One candidate, known format: read and forwarded untested, same as always — nothing extra needed.
- More than one candidate, known format: still forwarded untested, as
a list. The browser's
tryInterpreterCandidates(structile.html) tries each in order client-side — the same try-in-order-and-report contract asselect_interpreter, just run where the result will actually be used.widget'sinterpreter_sourcetrait isUnion[str, List[str]]for exactly this reason (seewidget.py). - Format itself unknown (a literal string, no path, no
format=) with aninterpreter=dict spanning more than one extension: forwarded asinterpreter_candidates_by_format({"xml": [...], "ini": [...]}— any mix of markup and other formats,formatleft""). The browser'stryInterpreterCandidatesAnyFormattries every entry (each under its own contract — DOM mode for markup, raw-text otherwise) and reports which format actually won — Python never runs anything just to decide that either.select_interpreter_any_format(used directly, or viaconvert()) does the equivalent server-side, needing themini-racerpackage either way since determining which of several formats applies is inherently a "parse it to find out" question.
The optional mini-racer package only re-enters the picture for
select_interpreter/select_interpreter_any_format used directly (not
through open()) and for convert(), below — both need a real
Python value back, which only actually running the interpreter (in an
embedded V8 engine — a real JS runtime, via a prebuilt wheel; no Node.js, no
npm, no system install) can produce. mini-racer is for
development/conversion work, never a requirement for viewing or diffing
data.
The "fail loudly" contract: a candidate "loses" the moment it raises —
either because the interpreter script itself throws, or because the DOM-lite
document couldn't be built at all (malformed XML — xml.etree.ElementTree
raises ParseError, which surfaces as a clean RuntimeError; mirrors the
<parsererror> check structile.html's own runInterpreter does
against a real browser DOMParser). HTML5 parsing never rejects anything
(it's error-tolerant by design, and structile's own html.parser-based
tree builder mirrors that leniency), so an interpreter meant for real HTML
data has to validate the document shape itself and throw if it's wrong —
see interpreters/generic_xml.js
(requires a root <o> element) and
interpreters/html.js (requires a real
<!doctype html>) for the two bundled examples' exact checks, and write
your own interpreters with the same expectation: prefer throwing over
silently returning something wrong when the input isn't actually your
schema. An interpreter that never throws can never lose a selection, even
against the wrong data.
interpreters/keyed_xml.js is a third,
mutually-incompatible example schema alongside generic_xml.js/
tagged_xml.js (see the "Custom formats via an interpreter" walkthrough in
python-library.md for the
service-config example all three share) — a distinct attribute-carrying
leaf shape (<item key="..." kind="..." value="..."/>) and a container
naming split (leaves use key=, containers use name= — see the file's
own header comment) that keeps the candidate-list trial mechanism honest
with three genuinely different schemas in play, not just two.
Interpreter pairing in the standalone viewer
Opening structile.html directly (no Python) and dragging/
picking a data file has no registry to consult, so it can't know in
advance which extensions might need an interpreter. classifyDroppedName
resolves this with a small, deliberately fixed whitelist —
STRUCTURED_EXT_FORMAT — rather than either hardcoding just xml/html/htm
(too narrow) or treating every unrecognized extension as "needs pairing"
(a behavior change for every currently-unhandled file, since a plain data
file under an unusual name/extension would stop getting its JSON/
Python-repr fallback attempt): .xml/.html/.htm (DOM contract) plus
.ini/.toml/.yaml/.yml (raw-text contract) and
.csv/.tsv/.tab (parsed natively — on the list so the extension
resolves to format "csv", but it only reaches the pairing flow if that
parse fails, or if an interpreter for "csv" is already loaded) — an
extension not on this list keeps the pre-existing behavior (attempt JSON/
Python-repr, error clearly if neither matches) unchanged.
For a file on the list, either half — the data file, or its .js
interpreter — can arrive first (paste, drag-drop, or a file picker), in
either order; pendingStructured holds whichever one showed up until the
other completes the pair (handleStructuredText/handleInterpreterText),
the same mechanism structile-embed-init's own pairing (a text+format with no
interpreterSource yet, waiting for a later structile-embed-interpreter
message — see above) reuses. Pasting itself (directly into the source pane —
see disableSourceView above) only ever recognizes XML-shaped raw text
(looksLikeXml — there's no filename to sniff a format from pasted text), so
pairing a non-markup format's data side always goes through a file, never a
paste.
This whitelist is a standalone-viewer-only concern — it doesn't exist on
the Python side at all. st.open(path, interpreter=..., format=...)
never needs it: an explicit interpreter=/registered extension is enough
regardless of what the extension is (see Custom formats via an
interpreter), since
Python always knows explicitly what's being asked for — there's no
"unrecognized file dropped with no other context" case to resolve there.
Distributing interpreters as a package
Writing and distributing interpreters covers the
end-user-facing mechanics (the pyproject.toml/register(registry) shape,
a copy-pasteable example) — this is the implementation picture:
structile/_plugins.py, driven from structile/interpreters.py's
get_registered_interpreter (the one function every resolution path —
open()/diff()'s markup-mode detection, resolve_candidates's own
registry fallback, convert()'s format inference — already goes through to
consult the registry).
Precedence, most to least specific:
| Rung | Source | Notes |
|---|---|---|
| 1 | An explicit interpreter= argument |
Always wins — never even reaches the registry. |
| 2 | A direct register_interpreter() call |
Wins regardless of when it happens relative to plugin discovery — see below. |
| 3 | A plugin registration (entry points) | Only reached once nothing above claimed the extension. |
| 4 | Nothing | Existing fallback behavior, unchanged (empty candidate list). |
Rung 2 beating rung 3 has to hold regardless of call order, which is
the one subtlety here: register_interpreter()/unregister_interpreter()
forget any plugin "ownership" of that extension the moment they're called
directly (see interpreters.py), and _plugins.py's registry facade
refuses to let a plugin overwrite an extension that's already present and
NOT plugin-owned. So a direct registration wins whether it happens before
first discovery (the plugin's own registration attempt is skipped, logged
at DEBUG) or after (a plain dict assignment overwrites the plugin's entry,
and a later load_plugins(force=True) won't silently reclaim it).
Discovery timing and caching: lazy — the module-level _loaded guard
flag in _plugins.py starts False and is never touched by import
structile itself, only by the first call that actually consults the
registry (get_registered_interpreter). Once set, discovery never runs
again in that process unless load_plugins(force=True) is called
explicitly (structile.load_plugins(force=True) — for a notebook session
where a plugin package was just pip installed and a kernel restart isn't
wanted). Each entry point is loaded and registered independently, each
wrapped in its own try/except Exception (ep.load() and the register()
call separately) — logged at WARNING with exc_info=True and never
applied, so one broken plugin never breaks another plugin's registration or
anyone's open()/diff()/convert() call. It's still recorded as a
failed PluginRecord (see "Introspection" below) rather than dropped
outright.
STRUCTILE_DISABLE_PLUGINS: set to any truthy value (anything other
than unset/empty/"0"/"false"/"no") to skip discovery entirely — logged
at DEBUG, not WARNING, since this is a deliberate opt-out, not a failure.
Existing direct register_interpreter() registrations are unaffected;
only entry-point scanning is skipped.
Collisions: if two plugins register the same extension, the second one loaded wins (entry-point iteration order isn't Python-version-guaranteed, so which one that is isn't predictable across environments) and a WARNING names both distributions — emitted regardless of which one happens to win, since the collision itself is the actionable information, not the winner.
Introspection: structile.plugins() returns one PluginRecord per
discovered entry point — successful or failed (entry_point,
distribution, version, extensions, plus success/error).
extensions is the list actually registered, which can be shorter than
what the plugin attempted if a rung-2 registration or another plugin
already claimed some of them; for a failed plugin (success=False) it's
always empty, and error holds a short "ExceptionType: message" summary
of whatever ep.load()/register() raised (the same exception already
logged at WARNING, just without needing logging turned on to see it).
Failures are included deliberately, not just successes — the person most
likely to reach for this is the one whose plugin isn't working, and for
them an entry silently missing from the list is indistinguishable from
"never installed". Triggers discovery itself, so it's safe to call as the
very first thing. python -m structile --plugins prints the same records
from the command line, rendering a failed one as name (dist version):
FAILED — error instead of its extension list.
InterpreterSource through both JS execution paths: an
InterpreterSource(text, name) is source text handed directly rather than
read from a path — the constructor argument a plugin's register() passes
after loading its bundled .js via importlib.resources. Every interpreter
candidate, however it arrives, is normalized to plain text in exactly one
place — structile/_paths.py's read_interpreter() — before it goes
anywhere else, so both JS execution paths described under "Interpreter
selection" above already handle it with no extra code: the browser-side
trial (open()/diff()) never even sees the InterpreterSource object,
only the .text it was read into by the time a Payload is built; the
mini-racer path (convert(), select_interpreter) reads it the same
way. The name only resurfaces in a log line if that particular candidate
loses a trial (select_interpreter's own warning logs %r of the
candidate object itself, before it's reduced to source text — an
InterpreterSource's repr() shows its name, never the source body). A
plugin's register() that hands _RegistryFacade.register_interpreter()
an InterpreterSource still carrying the default name="<inline>" gets a
DEBUG breadcrumb naming the distribution and extension — three candidates
all logging as <inline> in a failed-trial warning tells nobody which
plugin to blame, so a plugin author should always pass a real name=
(README's/tests/fixtures/structile_demo_plugin/'s examples already do).
convert()'s built-in/JS-runtime boundary
"json" <-> "python" is parsed and serialized entirely in Python (via
the stdlib json/ast modules — see convert.py's _parse_python_repr/
_serialize_python_repr) — no extra dependency, no subprocess, works
everywhere structile does. The moment either side is anything else
("xml"/"html", or a custom format like "ini"), an interpreter becomes
required (explicit interpreter=, or something registered for the
relevant extension — _infer_source's own extension-inference fallback
mirrors open()'s exactly: an unrecognized extension only becomes a
real format once something's actually attached to it) and runs as actual
JS in an embedded V8 engine (interpreters.run_interpreter_source, via the
optional mini-racer package) — exactly the way the browser would, since
an interpreter script is JS and reimplementing one in Python would mean
maintaining two copies of the same logic that could silently drift apart.
This needs pip install mini-racer — a prebuilt wheel bundling a real
V8, so no Node.js, no npm, no system install either way, for
"xml"/"html" or any other format. "xml"/"html" additionally need a
DOM to hand interpretXML/serializeXML, built directly from a
Python-parsed tree (dom_lite.py's stdlib ElementTree/html.parser —
not jsdom, not a general DOM implementation, just the subset
tagName/getAttribute/children/textContent/documentElement/
doctype/body every bundled interpreter and the documented contract
actually use — see dom_lite.js's header comment for exactly what that
means for a custom interpreter reaching for more); the raw-text contract
(interpretText/serializeText, for any other format) never touches a
DOM at all, so converting e.g. "ini" <-> "json" needs mini-racer
itself but never builds a DOM-lite document — proven by this project's own test suite
(a dedicated test monkeypatches the DOM-building function to fail loudly if
it's ever called on this path), not just claimed. Both directions are
optional capabilities, not a structile dependency; a clear
RuntimeError explains what's missing if you hit either path without
mini-racer installed.
Python's set/frozenset round-trip through every format this way (JSON
has no Set, so dst_format="json" flattens one to a sorted list, same as
open()'s own JSON encoding does); very large integers are exact
through "json"/"python" but, like everywhere else data passes through
an interpreter script, only as precise as that script's own parsing (the
bundled examples use parseInt/parseFloat, so expect ordinary
floating-point precision there, not exact big-int fidelity).
How a Set survives the Python<->JS boundary specifically: JSON (the
only shape that boundary can take, both for the value going into the
embedded JS engine and for the interpreter's own result coming back out)
has no Set type, so a Python value handed to run_interpreter_source for
serializing (interpreters.py's _mark_sets) has every set/frozenset
wrapped as {"__structile_set__": [...]} first — the one shape an
interpreter like interpreters/tagged_xml.js
(which has a dedicated <set> element, distinct from <dict>/<table>)
can recognize as "this array was really a set," not a list, once
dom_lite.js's own __structileUnmarkSets reconstructs it as a real JS Set on
the way in; the reverse (__structileMarkSets/_unmark_sets) handles a JS Set
an interpreter constructs coming back out of a "parse" run. This only
matters if you're writing your own interpreter or debugging why a Set
round-tripped as a list; every bundled interpreter and open()
itself (which never runs any JS for its own rendering — see "Interpreter
selection" above) already handle it correctly without you doing anything.
Config resolution, in full
Two independent stores share the same Options object and the same
three-tier precedence shape, but resolve completely separately and never
interfere with each other:
- Viewer settings (
gap,theme, theSETTINGS_SCHEMAlayout knobs) — resolved byconfig.resolve_config(overrides, extra)into the{settings?, theme?}object the viewer's embedded-mode protocol (and the standalone snapshot'sapplyInitialPayload) both expect. renderer— resolved separately byconfig.resolve_renderer(explicit, extra), and explicitly excluded from the objectresolve_configproduces (config._NON_VIEWER_OPTIONS) — it's Python-side dispatch only, the viewer has no concept of it at all.
Precedence for (1): per-call keyword (st.open(data, gap=8)) > per-call
config=Options() > module-level default (set_option/options.gap = 8)
the viewer's own built-in default (nothing sent at all). Precedence for (2): per-call
renderer=>STRUCTILE_RENDERERenv var > per-callconfig=Options()> module-level default (use()/set_option("renderer", ...)) auto-detect.
Command line reference
python -m structile data.json
python -m structile data.xml --interpreter interpreters/generic_xml.js
python -m structile data.json --renderer file --out snapshot.html
python -m structile data.json --no-open # write, don't open a browser tab
python -m structile --version
Same renderer resolution/auto-detection as calling open() from a
script (see "Renderers: how they actually work" above) — from an actual
terminal this almost always means browser. There's no --format/registry
flag on the CLI itself: --interpreter is inferred from the file's own
extension the same way open() infers it, and the registry
(register_interpreter) is a Python-API-only concept — there's nothing to
"register" for a single one-off CLI invocation.