High-performance Python bindings for the ZXC asymmetric compressor, optimized for fast decompression.
Designed for Write Once, Read Many workloads like ML datasets, game assets, and caches.
- Blazing fast decompression - ZXC is specifically optimized for read-heavy workloads.
- Buffer protocol support - works with
bytes,bytearray,memoryview, and even NumPy arrays. - Releases the GIL during compression/decompression - true parallelism with Python threads.
- Stream helpers - compress/decompress file-like objects.
git clone https://github.com/hellobertrand/zxc.git
cd zxc/wrappers/python
python -m venv .venv
source .venv/bin/activate
pip install .import zxc
data = b"hello zxc" * 1000
blob = zxc.compress(data) # level, checksum and dict are optional
assert zxc.decompress(blob) == datacompress() and decompress() accept any object supporting the buffer
protocol - bytes, bytearray, memoryview, NumPy arrays - and release the
GIL, so Python threads run them in true parallel.
zxc.compress(data, level=zxc.LEVEL_FASTEST) # fastest, largest output
zxc.compress(data, level=zxc.LEVEL_DEFAULT) # the default
zxc.compress(data, level=zxc.LEVEL_ULTRA) # smallest output, slowestLEVEL_FAST, LEVEL_BALANCED, LEVEL_COMPACT and LEVEL_DENSITY sit between
them. min_level(), max_level() and default_level() return the bounds at
runtime.
stream_compress() and stream_decompress() work on objects exposing a real
file descriptor via fileno(), so in-memory streams such as io.BytesIO are
not accepted. Both return the number of bytes written and take n_threads=0
to mean the default, auto-detected thread count.
with open("input.bin", "rb") as src, open("output.zxc", "wb") as dst:
written = zxc.stream_compress(src, dst, level=zxc.LEVEL_DEFAULT)
with open("output.zxc", "rb") as src, open("restored.bin", "wb") as dst:
zxc.stream_decompress(src, dst)For in-memory streams, use the push-based CStream / DStream instead:
with zxc.CStream(level=zxc.LEVEL_DEFAULT) as cs:
out = cs.compress(chunk) + cs.end()ZxcReader and ZxcWriter are io.RawIOBase adapters, so they compose with
io.BufferedReader and the rest of the io stack:
with open("output.zxc", "wb") as f, zxc.ZxcWriter(f) as w:
w.write(data)
with open("output.zxc", "rb") as f, zxc.ZxcReader(f) as r:
restored = r.read()detect_zxc(data) checks only the 4-byte magic word at the start of the
buffer. It does not validate the rest of the header, the footer, or the
archive contents, so a truncated or corrupt archive still passes - treat it as
content-type sniffing, not as a validity check.
A seekable archive can be decompressed block by block, without reading what comes before:
blob = ... # produced with stream_compress(..., seekable=True)
with zxc.Seekable(blob) as s:
print(s.num_blocks, s.decompressed_size)
middle = s.decompress_range(offset=1 << 20, length=4096)Cctx and Dctx carve their working buffers once instead of once per call,
which pays off when compressing many payloads with the same settings. A
dictionary given at construction applies to every call.
import zxc
with zxc.Cctx(level=zxc.LEVEL_DEFAULT) as cctx:
archives = [cctx.compress(p) for p in payloads]
with zxc.Dctx() as dctx:
payloads = [dctx.decompress(a) for a in archives]With a dictionary, the decoder must be given the same one:
with zxc.Cctx(dict=dictionary, dict_huf=table) as cctx:
archive = cctx.compress(payload)
with zxc.Dctx(dict=dictionary, dict_huf=table) as dctx:
assert dctx.decompress(archive) == payloadA context is not thread-safe: use one per thread.
Dictionaries lift the ratio on many small, similar payloads. The same dictionary is required at both ends:
d = zxc.Dictionary.train(samples) # samples: list[bytes]
blob = zxc.compress(data, dict=d.content, dict_huf=d.huf)
back = zxc.decompress(blob, dict=d.content, dict_huf=d.huf)
with open("corpus.zxd", "wb") as f: # persist
f.write(d.save())
with open("corpus.zxd", "rb") as f:
d2 = zxc.Dictionary.load(f.read())get_dict_id(archive) returns the dictionary id an archive was built with, so
you can pick the right one before decompressing.
cd zxc/wrappers/python
pip install pytest
pytest tests/ -vBSD-3-Clause - see LICENSE.