Skip to content

Latest commit

 

History

History
160 lines (114 loc) · 4.64 KB

File metadata and controls

160 lines (114 loc) · 4.64 KB

ZXC Python Bindings

High-performance Python bindings for the ZXC asymmetric compressor, optimized for fast decompression.
Designed for Write Once, Read Many workloads like ML datasets, game assets, and caches.

Features

  • Blazing fast decompression - ZXC is specifically optimized for read-heavy workloads.
  • Buffer protocol support - works with bytes, bytearray, memoryview, and even NumPy arrays.
  • Releases the GIL during compression/decompression - true parallelism with Python threads.
  • Stream helpers - compress/decompress file-like objects.

Installation (from source)

git clone https://github.com/hellobertrand/zxc.git
cd zxc/wrappers/python
python -m venv .venv
source .venv/bin/activate 
pip install .

Quick Start

import zxc

data = b"hello zxc" * 1000

blob = zxc.compress(data)                 # level, checksum and dict are optional
assert zxc.decompress(blob) == data

compress() and decompress() accept any object supporting the buffer protocol - bytes, bytearray, memoryview, NumPy arrays - and release the GIL, so Python threads run them in true parallel.

Compression Levels

zxc.compress(data, level=zxc.LEVEL_FASTEST)   # fastest, largest output
zxc.compress(data, level=zxc.LEVEL_DEFAULT)   # the default
zxc.compress(data, level=zxc.LEVEL_ULTRA)     # smallest output, slowest

LEVEL_FAST, LEVEL_BALANCED, LEVEL_COMPACT and LEVEL_DENSITY sit between them. min_level(), max_level() and default_level() return the bounds at runtime.

Streaming Files

stream_compress() and stream_decompress() work on objects exposing a real file descriptor via fileno(), so in-memory streams such as io.BytesIO are not accepted. Both return the number of bytes written and take n_threads=0 to mean the default, auto-detected thread count.

with open("input.bin", "rb") as src, open("output.zxc", "wb") as dst:
    written = zxc.stream_compress(src, dst, level=zxc.LEVEL_DEFAULT)

with open("output.zxc", "rb") as src, open("restored.bin", "wb") as dst:
    zxc.stream_decompress(src, dst)

For in-memory streams, use the push-based CStream / DStream instead:

with zxc.CStream(level=zxc.LEVEL_DEFAULT) as cs:
    out = cs.compress(chunk) + cs.end()

File Objects

ZxcReader and ZxcWriter are io.RawIOBase adapters, so they compose with io.BufferedReader and the rest of the io stack:

with open("output.zxc", "wb") as f, zxc.ZxcWriter(f) as w:
    w.write(data)

with open("output.zxc", "rb") as f, zxc.ZxcReader(f) as r:
    restored = r.read()

detect_zxc(data) checks only the 4-byte magic word at the start of the buffer. It does not validate the rest of the header, the footer, or the archive contents, so a truncated or corrupt archive still passes - treat it as content-type sniffing, not as a validity check.

Random Access

A seekable archive can be decompressed block by block, without reading what comes before:

blob = ...  # produced with stream_compress(..., seekable=True)

with zxc.Seekable(blob) as s:
    print(s.num_blocks, s.decompressed_size)
    middle = s.decompress_range(offset=1 << 20, length=4096)

Reusable Contexts

Cctx and Dctx carve their working buffers once instead of once per call, which pays off when compressing many payloads with the same settings. A dictionary given at construction applies to every call.

import zxc

with zxc.Cctx(level=zxc.LEVEL_DEFAULT) as cctx:
    archives = [cctx.compress(p) for p in payloads]

with zxc.Dctx() as dctx:
    payloads = [dctx.decompress(a) for a in archives]

With a dictionary, the decoder must be given the same one:

with zxc.Cctx(dict=dictionary, dict_huf=table) as cctx:
    archive = cctx.compress(payload)
with zxc.Dctx(dict=dictionary, dict_huf=table) as dctx:
    assert dctx.decompress(archive) == payload

A context is not thread-safe: use one per thread.

Pre-Trained Dictionaries

Dictionaries lift the ratio on many small, similar payloads. The same dictionary is required at both ends:

d = zxc.Dictionary.train(samples)          # samples: list[bytes]
blob = zxc.compress(data, dict=d.content, dict_huf=d.huf)
back = zxc.decompress(blob, dict=d.content, dict_huf=d.huf)

with open("corpus.zxd", "wb") as f:        # persist
    f.write(d.save())

with open("corpus.zxd", "rb") as f:
    d2 = zxc.Dictionary.load(f.read())

get_dict_id(archive) returns the dictionary id an archive was built with, so you can pick the right one before decompressing.

Testing

cd zxc/wrappers/python
pip install pytest
pytest tests/ -v

License

BSD-3-Clause - see LICENSE.