- Spec v3.5.1 conformance (SPEC 5, score-rounding errata). SPEC 5 now pins the graph
scoretwo-decimal wire form to round-half-to-even on the exact IEEE-754 double, resolving a midpoint divergence in the JavaScript and Kotlin SDKs. This SDK'sf"{score:.2f}"formatter already rounds half-to-even, so there is no behavior change; re-verified against the newgraph-encode/004_score_midpoint_roundingconformance fixture.
- Decoders now reject a declared
[N]section count that does not match the actual item count, in both directions, per SPEC Section 13 (Count Validation). A declared count smaller than the rows or entries present was previously read as a limit and the surplus was dropped; it is now an error. Covers the generic tabular, keyed-map, and root-array forms, the delta and full-set decoders, and the graph## edges [N]section. Valid payloads and encoding are unchanged.
- Keyed-tabular map encoding (SPEC 7.2a): a JSON object whose values are all objects forming a tabular set is encoded as a keyed table (
## [N:]{key,...}) - the shared value fields are declared once, with one key-prefixed row per member. Canonical by default, supported in nested and streaming positions, and integrated with generic delta using the map key as the identity.
- Negative zero is canonicalized to
0for both integer and floating-point values (SPEC 2.3.1). - Canonical-output alignment across all six SDKs: object key ordering, graph header fields, and symbol ordering follow the specification and reference implementation exactly.
- Conformance runners assert re-encode idempotence (
encode(decode(x)) == x) for the generic, graph, and delta profiles; a differential cross-SDK fuzz was added to the verification suite.
- The conformance runner now hard-fails on any unhandled operation (instead of silently skipping it) and exercises session, delta, roundtrip, and pack-root fixtures end to end; the graph delta wire decode and verify path is now covered, so no operations remain allow-listed.
- Implemented the graph delta wire decoder and verifier (
decode_delta/verify_delta): parse aGCF profile=graph delta=truewire back into removed/added symbols and edge changes, apply them atomically to a base snapshot, recomputepack_root, and reject a wrongnew_rootwithroot_mismatch(SPEC 10.4). The## addedencoder now emits the trailingdistancefield (SPEC 3.4.1, Section 10.1). The sharedgraph-deltafixtures now run end to end: 001 (encode, gains the trailing distance), 002 (verified apply), 003 (root_mismatchrejection). - Session encoding correctness fix.
encode_with_sessionassigned per-response local IDs instead of stable session-global IDs, so the cross-call dedup references (@N # previously transmitted) pointed at the wrong symbols, and the header emitted zero-valuedbudget/tokens/edges. Both are fixed to match the reference; graph session output is now byte-identical across all six SDKs. This had gone undetected because the conformance runner skipped the shared graph-session fixtures (now wired). - Added the graph-profile PackRoot (
pack_root(symbols, edges), gcf-pack-root-v1, SPEC 10.2): the content-addressed sha256 over canonical, independently-sorted symbol/edge records, byte-identical to gcf-go/rust/typescript/swift/kotlin. The conformance runner now exercises the sharedgraph-pack-rootfixtures, which it had been skipping (so this primitive was previously unimplemented and untested). - Buffered graph encoder now matches the reference byte-for-byte: symbols are ordered by distance then descending score with local IDs assigned in output order, and the header omits
budget/tokens/edgeswhen zero (previously symbols kept input order and zero-valued header fields were always emitted). The conformance runner now exercises the sharedgraph-encodefixtures (001-003), which it had been skipping - which is how this divergence went uncaught. - Buffered graph encoder: order edges by source ID, then target ID, then edge type (SPEC 16.1), instead of emitting them in input order. Decode-invariant (edges are a set) and does not affect
pack_root(which sorts edge records independently), so no content addresses change. Pinned by shared fixturegraph-encode/003. Streaming edges remain in producer-arrival order. - Decoder: reject an orphan
.fieldattachment (a.fieldwhose name is neither a^-marked column of its row nor a>-containing field name, SPEC 7.4.6.1.4) instead of silently absorbing it as an undeclared extra field. Such a stray attachment previously decoded to a record no encoder produces, silently injecting a field onto the last-parsed row (a lossless round-trip hole); now rejected per SPEC 16.5 (orphan_attachment). - Decoder: reject an orphan positional inline body (a pipe-delimited line with no eligible
^{}attachment-marker cell) instead of silently dropping it. The object-body parser previously skipped any unrecognized line, so a stray positional body (e.g. a secondBob|b@t.comafter a row's one inline cell was filled) vanished with no error (silent data loss); now rejected per SPEC 16.5 (orphan_inline_attachment). - Graph streaming trailer: the edge count is now always the last
countsentry, even when the stream has no edges (positionalcounts=2,1,0; labeledcounts=…,edges:0). A zero-edge stream previously dropped it, violating the SPEC §8.4 / §8.4.1 rule that the edge count is always present and last (the invariant that keeps the positional form unambiguous). The graph trailer is decoder-ignored, so this changes producer output only.
- New
labeled_trailer_countskeyword onStreamEncoder. When set, the##! summarygraph streaming trailer emitscounts=in the labeled formlabel:countper group (e.g.counts=targets:2,related:1,edges:3) instead of the default positional values-only form (counts=2,1,3). Default false is byte-identical to prior output. - Opt-in and non-breaking: a producer-side comprehension aid for known weak consumers. The trailer counts remain informational (decoder-ignored) in both forms; neither changes the decoded payload. Mirrors the
gcf-goreference.
- Streaming graph trailer now emits
distance_Ngroup counts in pure group-header emission order (dropping a fixedtargets,related,extendedprefix), matching the other SDKs and deterministic per SPEC 16.1. Byte-identical for contract-conformant (ascending-distance) input; pinned by shared fixturesstreaming-v2/010–011. - The conformance runner now executes the
graph-stream-encodefixtures (streaming-encode parity, previously decode-only): fixture 004 (positional trailer) and 005 (labeled trailer). - README: corrected the streaming example trailer from the defunct
## _summary … sections=to the real##! summary … counts=; README now leads with the project diagram. - Added a generic-delta fuzz test (decoder never crashes; string round-trip).
- Full producer + consumer implementation of generic-profile delta, byte-for-byte interoperable with
gcf-go:GenericSet(keyed record set),GenericDeltaPayloadgeneric_pack_root(gcf-pack-root-v1, generic profile) with a purpose-built cell canonicalization decoupled from the wire cell encoder: collision-free (null/bool/number bare, strings always quoted) and record-safe. Fields and records sort by UTF-8 byte order to match Go'ssort.Strings.diff_generic_sets(the blessed producer path; centralizes the keyed-diff invariants),encode_generic_full,encode_generic_deltadecode_generic_full,decode_generic_delta(consumer wire parsing)verify_generic_delta(atomic apply +new_rootverification)- Re-anchor session helper (SPEC §10a.8):
GenericDeltaSession(current_full,next) withReanchorPolicy/fixed_n(n)/size_guard()cadence policies andDEFAULT_REANCHOR_N = 15. Producer-side sugar over the primitives; introduces no new wire syntax (every emission is exactlyencode_generic_full/encode_generic_deltaoutput). Re-anchor cadence is byte-for-byte identical togcf-go(size guard uses UTF-8 byte length to match Go'slen(string)), verified by the sharedgeneric-delta-sessionconformance fixtures.
- Delta is opt-in and bilateral; the existing
encode_genericpath is unchanged (backward compatible).
- Unit suite mirroring
gcf-go: self-proving round-trip (diff -> encode -> apply -> recomputed root), determinism / row-order invariance, no-type-collision canonicalization, every invariant/error path, full-payload wire round-trip, the complete server -> wire -> consumer end-to-end loop, and malformed-wire-fails-closed. - Conformance runner support for
generic-pack-root,generic-delta,generic-delta-verify,generic-delta-decode(12 shared fixtures); verified to produce identical pack roots and delta wire togcf-go. - Session helper suite (
test_generic_delta_session.py) mirroringgcf-go: FixedN cadence pattern, size-guard triggering, schema-change forced full, FixedN(15)-over-30-turns count, and the load-bearing consumer-stays-in-sync check under both policies. Conformance runner support forgeneric-delta-session(3 shared fixtures: fixed-N, size-guard, schema-change). - Generic-delta fuzz (
test_generic_delta_fuzz.py), mirroringgcf-go: the decoder never crashes on arbitrary/mutated input, and arbitrary UTF-8 string cells (including multi-byte and control characters) survive the full-wire round-trip with the pack root preserved.
- Losslessness (nested null): a nested object that is null at an intermediate level (e.g.
{"meta": {"owner": None}}) is no longer flattened. Previously its leaves encoded as absent (~) and unflattened to a missing key, silently dropping the null. Such fields now fall back to the attachment mechanism; a top-levelNonestill flattens losslessly (emits-, reconstructs via the all-null rule). Enforced by the shared conformance fixturesflatten/017–019. Prototype pollution does not affect Python (dicts have no mutable prototype).
test_flatten_roundtrip: aligned arrays whose shared fields are fixed-shape nested objects with a field or an intermediate nested level sometimes null/absent — the shape the prior scalar-only generator never produced, leaving the flatten/unflatten path unexercised. Verified to fail on the pre-fix encoder and pass on the fix.
- Added
GenericOptionsdataclass withno_flattenfield to disable nested object flattening encode_generic(data, GenericOptions(no_flatten=True))produces attachment syntax instead of path columns- Backward compatible:
encode_generic(data)behavior unchanged (flatten on by default) - Fixed: field names containing
>no longer appear as tabular columns (spec rule 7.4.6.1.4) - Fixed: field names containing
>no longer eligible for flattening analysis - Fixed: decoder no longer treats literal
>in key names as a path separator - Fixed: decoder accepts orphan attachments (fields excluded from column list)
- 12 targeted edge case tests for
>in field names
- Encoder automatically flattens fixed-shape nested objects into
>path column names (e.g.,"customer>name"instead of^+.customer {}attachment) - Decoder reconstructs nested objects from
>path columns - 20-48% fewer tokens on deeply nested API data (Jira, Stripe, K8s, calendar events)
- 100% comprehension on every frontier model (validated across 9 models, 7 providers)
- Zero regression on lossless round-trips (230 tests, conformance + property-based)
- Falls back to attachment mechanism for: variable-length arrays, objects with different keys across rows, objects with
>in key names, empty nested objects
toolfield in graph profile header is now optional (SHOULD be present for MCP, not required)
- Quote strings containing commas (conformance:
inline-schema/006_inline_with_quoted_values) - Decode v2-format indented attachments in tabular rows (conformance:
decode/002_attachment) - Reject duplicate attachments on the same row (conformance:
errors-v2/027_duplicate_attachment)
encode_genericnow produces inline schema format (not backwards compatible with v1.x decoders)- Attachment lines no longer indented (same depth as parent row)
- Inline object fields use positional encoding without field-name prefix
- Inline object schema: objects with 3+ scalar fields encoded positionally with
^{fields}header - Shared array schemas: identical nested arrays omit
{fields}after first row - 472M+ fuzz iterations across all 6 implementations, zero failures
- Quote strings starting with
.(dot prefix) - Quote C1 control characters (U+0080-U+009F)
- Quote Unicode whitespace (NBSP, hair space, etc.)
- CLI:
encode-genericanddecode-genericsubcommands for generic profile - CLI now supports both graph and generic profiles
python -m gcfentry point
SPEC v2.0 implementation. 126/133 conformance fixtures passing (7 skipped: session, delta, binary UTF-8, negative zero, graph encode). 40M property-based round-trips with zero failures.
encode_genericemitsGCF profile=genericheaderdecode_genericrequiresGCF profile=header- Strings colliding with typed literals are quoted
- Full JSON string escaping and number grammar
-for null,~for absent,^for nested attachments##! summarytrailer replaces## _summary- Graph encoder emits
profile=graph
scalar.py: common scalar grammar (quoting, escaping, parsing, number formatting)- Conformance test runner (133 fixtures)
- Property-based round-trip tests (40M verified, configurable via
GCF_ITERATIONS)
GenericStreamEncoder: zero-buffering tabular streaming encode (begin_array/write_row/end_array/write_kv/write_section/write_inline_array)decode_generic: decode any GCF text (tabular or graph) back to Python objectsStreamEncoder: zero-buffering streaming encode (added in v0.4.0)
encode_generic: primitive arrays inlined asname[N]: val1,val2,val3
- Breaking:
encode()now emitsedges=Nin header line - Breaking:
encode()now emits## edges [N]section header (was## edges) decode()updated to parse## edges [N]format (strips bracket suffix)- Session encoder updated to emit new edge count format
- Docs: update README for PyPI discoverability (gcformat.com, proxy, vs-toon links)
- Fix: decoder rejects headers missing required
toolfield (conformance) - Fix: escape newlines as
\nin quoted strings inencode_generic
- Fix: escape
"inside quoted strings inencode_generic - Fix: quote empty strings as
""per spec
encode_generic: encode arbitrary Python values into GCF tabular format- Tabular encoding: positional rows with pipe separators, section headers, nested field support
- Uniform array detection with 70% key overlap threshold
- Initial release
encode/decode: full GCF round-tripencode_with_session: session deduplication (92.7% savings by 5th call)encode_delta: delta encoding for re-queries (81.2% savings)- Thread-safe
Sessionclass - 16 kind abbreviations
- CLI:
gcf encode,gcf decode,gcf stats - Type hints, Python 3.9+, zero runtime dependencies