Skip to content

Choosing obj, format, and interpreter

A practical guide to structile.open() (and structile.diff(), which follows the exact same rules independently for each side) — for when the short version leaves a specific case unclear.

The two modes

Every call to open() ends up in one of two modes, decided by what you pass for the data itself (obj), format=, and interpreter= together.

Plain value mode (the default) — what you get with no format= and no interpreter=:

  • Pass an ordinary Python value — a dict, list, string, number, None, a set, whatever open() supports — and it's shown and edited as that value, JSON-style.
  • Pass a path to a file instead, and it's loaded from disk first. (A plain str only counts as a path when that file exists; otherwise it's just the string itself, so a mistyped filename shows up as a one-word document. A pathlib.Path is always a path, and a missing one raises FileNotFoundError.) If the file's contents happen to be valid JSON, its exact original text is what you see and edit (whitespace, key order, everything preserved until you actually change something). Otherwise, contents that are a Python literal — a .py data file like {'a': (1, 2), 'b': None} — are read as that value, just as the viewer reads them; anything else on disk is loaded and shown as plain text.
  • Except .csv/.tsv/.tab, which Structile parses itself: those are handed to the viewer as their own text with format="csv" and no interpreter — the third built-in format, not a special case of plain value mode. The delimiter is detected from the content, so both spellings share one format label.

Interpreter mode — what you get once any of the following is true: you pass interpreter=, you pass format= set to anything other than "json"/"python" (e.g. "xml", "ini", or a format of your own — format="csv" takes this branch too, but resolves to the built-in parser instead of requiring an interpreter), or you've already called register_interpreter() for the file's extension earlier in the program (or an installed interpreter plugin package did it for you). In this mode obj is never read as a Python value — it's handed to an interpreter script as raw text, which turns it into something structile can display (and, if the script supports it, turns edits back into text on save). See Writing and distributing interpreters for what an interpreter script is and how to write one.

There are two ways to supply that raw text: a path to an existing file (the format is inferred from its extension — .xml → "xml", .ini → "ini", .csv/.tsv → "csv", and so on), or the text itself, as a plain string — e.g. xml_text = "<a><b>1</b></a>" handed straight to open(), with nothing on disk. format= is required in this case: a string has no filename to read an extension from, and — this is the part that trips people up — passing interpreter= doesn't fill that gap either. An interpreter script is just code; nothing about it says which format it's meant to read ahead of time. interpreter= only tells open() that this is interpreter mode; format= is what says which format.

st.open(xml_text, format="xml", interpreter="interpreters/generic_xml.js")

diff() follows all of the above too, independently for its left and right values — one side can be plain value mode while the other is interpreter mode (e.g. diffing a JSON file against an XML one).

A common surprise

Opening a file with an unregistered extension — .xml, .html, .ini, anything except the built-in .json/.csv/.tsv/.tab — with no format= and no interpreter= never fails. It silently falls back to plain value mode and shows the file as a quoted block of text instead of parsed data. If a markup file you expected to render as a document instead shows up looking like one giant JSON string, this is almost always why. Fix it with either:

st.open("data.xml", interpreter="interpreters/generic_xml.js")   # this call only
st.register_interpreter(".xml", "interpreters/generic_xml.js")   # every later .xml file

The full combination matrix

For the exhaustive, orderable/filterable version of everything above — all 17 kinds of obj × 7 format values × 7 interpreter shapes, 833 rows, each one actually run against the real code, not hand-derived — see below, or open it fullscreen.