Description
ZarrParser fails to open Zarr v2 stores containing vlen-utf8 (or vlen-bytes/vlen-array) string arrays whose .zarray has no explicit fill_value (i.e. "fill_value": null), which is the default for how xarray/zarr write string variables:
ValueError: Zarr data type resolution from object failed. Attempted to resolve
a zarr data type from a numpy "Object" data type, which is ambiguous, as
multiple zarr data types can be represented by the numpy "Object" data type.
Root cause
In virtualizarr/parsers/zarr.py, _convert_v2_to_v3_dict synthesizes a default fill value when the v2 metadata omits one:
if metadata.fill_value is None:
v2_dict = metadata.to_dict()
v2_dtype = parse_dtype(cast(Any, v2_dict["dtype"]), zarr_format=2)
fill_value = v2_dtype.default_scalar()
...
parse_dtype(v2_dict["dtype"], zarr_format=2) is called with only the bare numpy dtype string ("|O" for an object-dtype array), which is inherently ambiguous — zarr-python can't tell whether that's vlen-utf8, vlen-bytes, or a vlen array without more context.
But metadata (the ArrayV2Metadata passed into this function) has already resolved this correctly, via filters (e.g. numcodecs.VLenUTF8), to VariableLengthUTF8:
>>> zarr.open(store=..., path="trajectory", mode="r").metadata
ArrayV2Metadata(..., dtype=VariableLengthUTF8(), filters=(VLenUTF8(),), ...)
The bug is that this line re-derives the dtype from scratch instead of reusing metadata.dtype, throwing away the disambiguating information ArrayV2Metadata.from_dict already extracted from filters.
Minimal reproduction
Fully offline, no cloud credentials required:
import numcodecs
import zarr
from obstore.store import MemoryStore
from zarr.core.dtype import VariableLengthUTF8
import virtualizarr as vz
from virtualizarr.parsers import ZarrParser
from virtualizarr.registry import ObjectStoreRegistry
BUCKET_URL = "memory://b"
STORE_URL = "memory://b/store.zarr"
store = MemoryStore()
zstore = zarr.storage.ObjectStore(store, read_only=False)
root = zarr.open_group(store=zstore, path="store.zarr", mode="w", zarr_format=2)
# a v2 vlen-utf8 string array with no explicit fill_value -- exactly how
# xarray/zarr write string variables by default
arr = root.create_array(
"trajectory",
shape=(3,),
chunks=(3,),
dtype=VariableLengthUTF8(),
filters=[numcodecs.VLenUTF8()],
compressors=None,
fill_value=None,
)
arr[:] = ["a", "bb", "ccc"]
registry = ObjectStoreRegistry({BUCKET_URL: store})
vds = vz.open_virtual_dataset(url=STORE_URL, registry=registry, parser=ZarrParser())
print(vds)
Traceback:
File ".../virtualizarr/parsers/zarr.py", line 294, in _convert_v2_to_v3_dict
v2_dtype = parse_dtype(cast(Any, v2_dict["dtype"]), zarr_format=2)
...
ValueError: Zarr data type resolution from object failed. Attempted to resolve
a zarr data type from a numpy "Object" data type, which is ambiguous, ...
Also reproduces against real public data
s3://sofar-spotter-archive/spotter_data_bulk_zarr (variable trajectory, dtype: "|O", filters: [{"id": "vlen-utf8"}], fill_value: null) fails identically.
Suggested fix
In _convert_v2_to_v3_dict, reuse metadata.dtype (already correctly resolved by ArrayV2Metadata.from_dict) instead of re-deriving it via parse_dtype(v2_dict["dtype"], zarr_format=2):
if metadata.fill_value is None:
v2_dict = metadata.to_dict()
v2_dtype = metadata.dtype # already resolved, handles vlen-utf8/vlen-bytes/etc.
fill_value = v2_dtype.default_scalar()
...
Environment
virtualizarr==2.7.1 (latest release)
zarr==3.2.1
numcodecs==0.16.5
- Python 3.13
Description
ZarrParserfails to open Zarr v2 stores containingvlen-utf8(orvlen-bytes/vlen-array) string arrays whose.zarrayhas no explicitfill_value(i.e."fill_value": null), which is the default for how xarray/zarr write string variables:Root cause
In
virtualizarr/parsers/zarr.py,_convert_v2_to_v3_dictsynthesizes a default fill value when the v2 metadata omits one:parse_dtype(v2_dict["dtype"], zarr_format=2)is called with only the bare numpy dtype string ("|O"for an object-dtype array), which is inherently ambiguous — zarr-python can't tell whether that'svlen-utf8,vlen-bytes, or a vlen array without more context.But
metadata(theArrayV2Metadatapassed into this function) has already resolved this correctly, viafilters(e.g.numcodecs.VLenUTF8), toVariableLengthUTF8:The bug is that this line re-derives the dtype from scratch instead of reusing
metadata.dtype, throwing away the disambiguating informationArrayV2Metadata.from_dictalready extracted fromfilters.Minimal reproduction
Fully offline, no cloud credentials required:
Traceback:
Also reproduces against real public data
s3://sofar-spotter-archive/spotter_data_bulk_zarr(variabletrajectory,dtype: "|O",filters: [{"id": "vlen-utf8"}],fill_value: null) fails identically.Suggested fix
In
_convert_v2_to_v3_dict, reusemetadata.dtype(already correctly resolved byArrayV2Metadata.from_dict) instead of re-deriving it viaparse_dtype(v2_dict["dtype"], zarr_format=2):Environment
virtualizarr==2.7.1(latest release)zarr==3.2.1numcodecs==0.16.5