Changelog¶
All notable changes to torchfits are documented here.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
Unreleased¶
1.1.3 — 2026-09-09¶
Fixed¶
FITSHeaderScale/FITSScaleColumnskeep integer inputs as float afterBSCALE*x+BZERO(the unsigned conventionBZERO=32768on int16 no longer wraps; fractional BSCALE is no longer truncated)._median/_quantilepromote float16/bfloat16 to float32 (torch.quantilerejects half dtypes);LogStretchupcasts before1 + a*xso float16 no longer overflows to inf.SigmaClip(dim=(), fill="median")fills with the global median instead of a per-element identity.FITSHeaderNormalizemaps integer counts through float64 when the storage width is 32 bits or more, so values above2**24stay distinct.table.read_torch(where=...)with nocolumns=returns every column (matching Arrowtable.read);table.schema(columns=[...])uses the requested field order.FITSFile::read_subsetreturns the image's true dtype for degenerate (zero-width/height) cutout boxes instead of always float32.HDUList.fromfileno longer reads every header twice (headers from the batch open are reused), and theHDUInfo.headerbinding preserves duplicate HISTORY/COMMENT cards.TableHDURef.replace_column_filemirrors the newTFORMinto the in-memory header (matchinginsert_column_file).Headerconstructed from 3-tuples normalizes keys and scalar values like every other set path.- HTTP Range HDU walks account for BINTABLE heap size (
PCOUNT/THEAP), so cutouts past VLA columns or compressed images no longer fall back to a full-file download when the heap exceeds the scan cap. - CLI
arith --hdu2reads operand B once instead of per A-HDU;arith/compress--out-dirnames strip CFITSIO[section]suffixes;statsupcasts integer images once. bench_arrow_tables.pyuses the currenttable.readsignature.
Changed¶
- Derived
TableHDUs (filter,select,head,add_column, ...) own a copy of the source header, so later mutations of the derived header no longer leak into the parent. - CLI
transformpreserves safe header keys (WCS, EXTNAME, etc.) on float outputs instead of dropping the whole header; only scaling keywords (BSCALE/BZERO/BLANK,DATAMIN/DATAMAX) and checksum stamps are dropped. - CLI
compressresolves IO pairs through the shared batch resolver. get_cache_statsno longer reports counters that were never updated.- Removed the unreachable
tzero is Nonebranch infits_schema.
Added¶
InterquantileScaletransform for robust scaling by interquantile range;FitsStagedCutoutIterableDatasetcan load companion HDUs (e.g. label or mask maps) alongside the image HDU (#241).
Docs¶
- Documentation rewritten in plain user-facing language: internal tracking codes and process jargon removed, benchmark and install pages glossed.
1.1.1 — 2026-08-29¶
Fixed¶
table.read_torch(where=)drops TNULL sentinels and uses range-safe compares, matching Arrowtable.read.- Robust-quantize BLANK pixels are NaN on
torchfits.read, including uncompressed scaled images. Native IEEE float/double HDUs keep Inf and signed zero (fits_read_imgnulval is only for BLANK and compressed tiles; CFITSIOfnanwould otherwise replace them). torchfits copyis a byte copy (shutil.copy2); same-path I/O is refused;HDUList.writeto an existing path uses tempfile+replace.TableHDURef.headcomposes an existing row window.- Header cache hits clone
Headerfrom cards;raw_scalereaches the fallback reader. replace_hdustrips staleZ*cards; checksum rewrites restamp when the input had stamps;verify_checksumsreportspresent.- Table
.fits.gz/.ziprefuse the buffered pread path; TSBYTE mmap matches CFITSIO; TFORM repeat overflow raises. - CLI: unsigned
diffmin/max, Ctrl-C exit 130, JSON without NaN. - HTTP Range cutouts of integer images with
BLANKfall back to CFITSIO so missing pixels are NaN rather than the sentinel code. - Writing an already-decoded float with a copied integer header drops
BSCALE/BZERO/BLANKso CFITSIO does not scale twice (CLI cutout and arith). - Quantized table
TNULLvalues are Arrow nulls, sowhere="V IS NULL"matches. Native float NaN withoutTNULLnis unchanged. - Import sets
KMP_DUPLICATE_LIB_OKwhen unset (macOS libomp). - build: Portable sha256 verification in vendor.sh (macOS runners)
- changelog: Rank final tags above prereleases in latest_tag
-
build: Select sha256 tooling by OS, not binary presence
-
Completed the 1.1 correctness review for silent NaN/TNULL/copy bugs (#237)
- Completed the 1.1.1 release-readiness review (19 tracked findings closed)
Changed¶
torchfits._cppno longer re-exports undocumented_Cnames.- Root
to_astropy(path)delegates totable.to_astropy. read(hdu=[...], mmap=False)and path-list batch honor mmap.quantize=onHDUList+compress=raises instead of being ignored.- Linux wheels install
bzip2-develin manylinux soHAS_BZIP2matches macOS. - Sanitizer CI uses the pixi
testenv (not uv) and quotes cmake define flags so bash does not split on;.
Docs¶
- Benchmark headline cites git-mirrored
exhaustive_*_20260807_013736. GPU copy is host-decode thendevice=. Complex columns are Partial on Arrow. Release runbook uses OIDC, not a PyPI token. Windows is unsupported.CacheConfig.max_filesis a no-op.
Added¶
- 1.1.1 identity-check test suite and example-gallery additions
1.1.0 — 2026-08-26¶
Feature + correctness release on the same 2.13 torch ABI lane. New
capabilities: checksum-stamped writes, GIL-free hot reads, clean
truncated-file errors, high-fidelity Astropy interop, memory-bounded
streaming filters, multiprocess-safe remote downloads, auto-adaptive
FITS RGB, and native whole-file .bz2 FITS reads — plus a
major-release-readiness review that cleared a round of silent-corruption,
thread-safety, and supply-chain fixes.
Features¶
-
Native whole-file
.bz2FITS reads on builds with bzip2 support (capability flagtorchfits._C.HAS_BZIP2; vendored CFITSIO links libbz2 from the conda prefix or the system): every reader entry point decompresses transparently, while direct-I/O fast paths automatically route through CFITSIO. Writing a.bz2-named output is rejected — CFITSIO would silently create an uncompressed file; usecompress="BZIP2_1"for tile compression instead. -
transforms.rgb(*bands)auto-adaptive RGB from 1–7 aligned filters (shortest wavelength first): scarlet-style mix, per-band sky-median subtract + MAD equalize unlesscalibrated=True/zeropoints=, fitspng-style MAD stretch, coupled asinh, saturation, and sRGB. NaN mosaic holes stay black (they are not treated as sky). Scene classification is blow-out-safe: a bright extended object (planet disk, galaxy core) whose post-equalize p90 towers above the noise is stretched against its own p90 instead of the faint-feature anchor, so extended targets never clip while deep/star fields keep their stretch.lupton_rgbstays the Astropy-parity 3-band mapping (reddest first). write(..., checksum=True)stamps CFITSIODATASUM/CHECKSUMkeywords on every HDU at write time (all payload types, compressed included); verify later withtorchfits.verify_checksums.- Multithreaded reads scale: hot C++ paths (
read_full,read_full_numpy, header/shape/HDU-count probes) now release the GIL for open + I/O, so DataLoader threads no longer serialize behind one Python thread during disk or network access. - Truncated files raise instead of crashing: mmap table reads, filtered scans, and row updates validate the header-claimed extent against the real file size and raise a clear "truncated" error rather than SIGBUS-killing the interpreter.
- Multi-chunk buffered reads are correct again: tables whose rows
span more than one 16 MiB scratch chunk (payload > ~16 MB with caching
disabled) could read from an unsized staging buffer after a prefetch
rewrite — slow, garbage results on wide/mixed projections. Buffer
rotation now happens only while prefetching; regression tests pin
16 MB reads against astropy.
- High-fidelity
table.to_astropy(): TNULL-bearing columns become realMaskedColumns (no more object-dtype degradation), TUNIT maps to.unit, and fixed-size vector columns keep their(N, repeat)shape. - Memory-bounded streaming filters:
scan(..., where=...)now evaluates predicates per batch as rows stream past — peak RAM tracksbatch_size, not table size. Hidden predicate columns are projected out after filtering, and a fully filtered-out scan still yields one typed empty batch. - Multiprocess-safe remote downloads: concurrent DataLoader workers
sharing a cache directory are serialized by an OS file lock (exactly
one fetch per URL), interrupted transfers resume via
If-Rangevalidators (stale partials restart cleanly instead of producing hybrid files), and servers withoutContent-Lengthtrigger an explicit completeness warning.
Changed¶
torchfits convertPNG default is auto RGB (--recipe auto): 1–7 files, blue→red order,--brightness/--saturation/--calibrated/--zeropoints.--recipe luptonkeeps the previous 3-band reddest-first--q/--stretchmapping.- Scalar-column shapes are now rank-1 everywhere: FITS repeat==1
columns read as
(N,)through every path —hdul[n].data[col],hdul[n][col],TableHDURef,read_torch,iter_rows,to_tensor_dict, streaming chunks, and buffered/mmap reads alike. Vector columns (repeat>1) keep their shapes; packed string columns are never squeezed. Completes the accessor-only squeeze shipped in #234. read_torch(start_row=..., where=...)now filters within the row window, matchingtable.read(row_slice=..., where=...)and the Arrow engine. Previously the two entry points returned different populations for the same query.- WHERE negation follows SQL three-valued logic on every engine:
NOT (X == 5),X NOT IN (...), andX NOT BETWEEN ...exclude NULL rows exactly like their unnegated counterparts. - TNULL sentinel values can no longer satisfy numeric predicates on the C++ pushdown / torch-mask paths; sentinel-only matches are excluded so all engines return identical rows for the same query.
-
Batch fast paths (
read(list_of_paths), list-of-HDUs) honorfp16/bf16/raw_scaleinstead of silently returning raw data. -
torchfits.cppis deprecated in favor of privatetorchfits._cpp; every attribute access warns. No-op native cache stubs (configure_cache,get_cache_size,clear_file_cache) left the public surface. TORCHFITS_CFITSIO_CACHE_MB/_FILESenv vars removed; they configured a native cache that no longer exists.cache.configure_cache()/CacheManager.configure_cpp_cache()remain as documented no-ops emitting DeprecationWarning for one cycle.write_tensor()acceptschecksum=for parity withwrite().torchfits.open()accepts onlymode="r"; in-place update modes are rejected with an actionable error.TensorHDU.chunks()is implemented: lazy row-band slabs equal to slices ofto_tensor(). It previously called a binding that never existed and always raised AttributeError.DataView.dtypereports convention dtypes (uint16/uint32/int8) matching reader output rather than raw storage BITPIX.- TableHDU schema caches hold strong header refs so GC id-reuse cannot
serve stale schemas;
select()rejects unknown columns andhead()validates its argument. - Vendored CFITSIO fetches are sha256-pinned and verified; unpinned tags
fail closed unless
TORCHFITS_VENDOR_ALLOW_UNPINNED=1; "latest" resolution removed. Conda/pixi builds compile the same pinned+patched CFITSIO as the wheels — PLIO buffer fix + BZIP2_1 on every channel.
Fixed¶
- BIT (
'X') table columns wrote corrupted bits whenrepeat % 8 != 0(e.g.'12X'): both the initial writer and the buffered row-update writer passed flat element runs tofits_write_col(TBIT), which maps them onto raw data-unit bits ignoring per-row padding. Every write call is now confined to a single row; verified byte-exact against astropy for 8X/12X/16X/23X across all three write paths. - Buffered table reads no longer leak a closed fd after a transient pread failure — a recycled descriptor could later serve an unrelated file's bytes as table data.
- Cached reads are isolated from caller mutation: in-place edits of a returned tensor no longer poison subsequent identical reads (the default read cache stores/hands out private copies).
replace_hduheader preservation drops staleBSCALE/BZERO/DATASUM/CHECKSUM; grafting an unsigned-convention BZERO onto new float data previously made every reader misinterpret the replacement.SigmaClip: a single NaN no longer wipes the whole frame (valid mask is seeded from finite values even without a user mask); median fill of fully-masked groups yields 0 instead of NaN.GlobalScalarNorm: negative statistics divide sign-preservingly instead of exploding to ~1e30 scales;inverse()round-trips exactly; NaN no longer poisonsmean/rmsstatistics.- WHERE string literals survive parsing:
NAME == 'AT&T'matchesAT&T(notAT AND T), FITS doubled-quote escapes work ('O''NEIL'), and unterminated quotes raise instead of mangling. - Header parser: LONGSTRN chains keep assembling correctly when any card
carries a comment, and ESO
HIERARCHcards parse to typed values under their full keyword (ESO TEL AMBI TEMP -> 12.5). - Thread-local HDU metadata cache rotates its generation id when an out-of-band file replacement is detected, so stale shape/dtype/scale cannot be paired with new bytes on the mmap path.
-
Unsigned-convention detection tolerates floating-point imprecision in stored
BZERO/TZEROuniformly for images and tables (a file whose offset was serialized as 32767.999… now reads as uint16/uint32 on every path). -
Compressed-image null pixels decode as NaN (was silent 0): the null probe targeted a CFITSIO symbol that exists in no upstream release. ZBLANK is probed directly and float CompImage reads always pass NaN nulval. Regression suite vs astropy: tests/test_compressed_nulls.py.
torchfits arithno longer wraps/truncates integer images: ops run in int64/float64 with saturating cast-back plus warning;--dtypeoverrides; div produces floats underauto.torchfits statsworks on unsigned-convention images: min/max run after upcast instead of raising on missing uint reduction kernels.quantize="robust"maps NaN/Inf to BLANK sentinel codes with the keyword written, instead of packing non-finite values into valid codes that dequantize as real data.- WHERE float equality is engine-independent; schema() stops
lying about complex columns; row windows keep VLA/string
columns aligned; scan(mmap=True) handles scaled tables and
ASCII HDUs; chunk buffers zero-initialized; read_batch
warnings document skip semantics; interop kwargs no longer leak to
pandas; FITS numeric parsing accepts D-exponents / rejects
1_0. - Transforms are functional (no caller-tensor mutation via
.to()aliasing); medians interpolate like numpy/astropy; SigmaClip/AsymmetricSigmaClip gainfill="nan". - Random Groups images fail loudly instead of decoding garbage.
torchfits difftreats NaN == NaN: byte-identical files containing NaN pixels compare clean instead of reporting spurious differences.torchfits --transform -Jfan-out builds one transform instance per worker file, so stateful transforms never share_last_stateacross threads.- Thread-safety: TensorHDU reads use private per-call handles; shared TableReader instances serialize I/O. Stale-cache windows after header mutations closed by invalidating SharedReadMeta/readers in the header-card/key/checksum writers.
- packaging: SPDX license expression for license-files; boundary test tracks _cpp move
- tables: Engine-aligned WHERE floats, honest complex schema, aligned windows, streaming fallbacks
- cli: Integer-safe arith, uint stats, NaN-aware diff, per-worker transforms
Security¶
read_hdusenforces the same SSRF/path guards as every other entry point (loopback/link-local/private targets were reachable before).- HTTP credentials (
TORCHFITS_HTTP_TOKEN/TORCHFITS_HTTP_AUTHORIZATION) are stripped when a redirect crosses origins; kept only for same-origin hops and plain http->https upgrades.
Performance¶
- Policy/meta caches (
image_meta, cold-nommap, auto-mmap, hdu-type) are validated against the file's stat signature, eliminating stale dispatch after in-place rewrites. Measured cost of the validation: ~1.6 us peros.staton this host. - Cached reads now hand out private copies, so callers can mutate results.
Measured on a 64 MiB float32 image (local NVMe): uncached read 46.5 ms
(~1.4 GB/s) vs warm cache hit 45.3 ms including the isolation copy — a
hit still beats the I/O it replaces, and repeated table hits stay
sub-millisecond. If a cached-hot workload regresses measurably for you,
read(..., cache_capacity=0)restores v1.0 semantics at the cost of re-reading. - Full CPU + CUDA exhaustive benchmark re-runs on Linux CANFAR headless (published
CSVs
exhaustive_cpu_20260807_013736/exhaustive_cuda_20260807_013736underdocs/assets/bench/; lab profile, mmap on+off matrix, 3057 + 4315 rows): 100% of significant image comparisons won on both hosts. Exactly one case family remains where a peer leads: narrow-table full reads withmmap=Falsetrail fitsio by 21-36% on CPU and 8% on CUDA (buffered path stages whole rows; single-pass decode lands in 1.2). Image HCOMPRESS lags vs fitsio are sub-1.03× noise. A double-buffered chunk prefetch now overlaps the buffered path's I/O with decode for payloads >= 64 MB (gated from an earlier 4 MB threshold after CANFAR A/B showed thread handoff regressing warm-cache 13 MB tables). Benchmark harness fairness fixes: device synchronization on GPU timings, seeded interleaving order, medians over means, and cache-symmetric peer comparisons. - Multi-HDU writes flush process-global caches once per operation instead of twice per HDU.
- BIT (
'X') writes now issue onefits_write_colcall per row; only tables containing bit columns pay for the extra calls.
Added¶
- Table mutations warn on silent value loss: float payloads into integer columns (truncation/non-finite counts), out-of-range integers, non-ASCII characters dropped from string columns, and over-width string clipping.
- io: Native whole-file .bz2 FITS reads on bzip2-capable builds
Dependencies¶
- Vendored CFITSIO updated to 4.7.0 for wheel, source, conda and pixi
builds alike (
extern/VERSIONS.txt, sha256-pinned): every channel ships the PLIO buffer fix and BZIP2_1 support (see architecture).
Docs¶
- Scalar-column shape contract documented; architecture note reconciled with the wheel-vs-conda CFITSIO split.
- Corrected inaccurate benchmark and compatibility claims in the docs and toned down unverifiable ones
- changelog: Versionless Unreleased + generator tooling; refresh roadmap
- Changelog entries curated under Unreleased; changelog-check green
- io: Document BZIP2_1 availability, lossiness and interop caveat in write()
1.0.0 — 2026-08-09¶
Version cut on the 2.13 torch ABI lane: buffered table reads through a thread-local reader cache (process-wide eviction, stat-identity stale guard), single-open insert/update of table rows, and torch-mask predicate materialization. Final pre-release checks complete; awaiting collaborator docs/usability testing before the tag.
Added¶
torchfits.to_astropy(): Direct conversion of PyTorch tensor dictionary structures to AstropyTableinstances via zero-copy Arrow intermediate buffers.- Streaming ML Datasets:
FitsCubeIterableDataset,FitsSpectrumIterableDataset, andFitsStagedCutoutIterableDatasetwith rank and world size sharding for distributed PyTorch training. - MegaCam cosmic-ray denoise example (
example_megacam_cr_denoise.py): Noise2Noise on real dark/bias calibration twins (zero-field N2N), with self-normalizing pair transforms, held-out CCD evaluation, and honest transfer metrics (CR suppression, star fluxes, background, noise-injection test). Full rationale and results: Denoise pipeline. scripts/fetch_cfht_calib_frames.sh: idempotent download of CFHT MegaCam darks/biases from the CADC data service.scripts/canfar_denoise_incontainer.sh+scripts/launch_canfar_denoise.sh: headless Skaha job running the denoise example on a CUDA GPU (defaults to the 4-epoch setting; longer fixed-LR runs diverge — see the pipeline page).- Denoise example writes a before/after gallery figure
(
examples/output/megacam_cr_denoise_dark.png, rendered in the ML guide and the pipeline page). - Exemplary science-pipeline benchmark (
bench_science_pipeline.py: robust sigma-clipped coadd + ML cutout serving, torchfits vs astropy) with per-stage code-lines rows (bench_contract.code_lines);bench-denoisepixi task for the MegaCam CR-cleaning benchmark.
Performance¶
- Cached table reader:
read_table/read_batchreuse an open CFITSIO handle per thread instead of open/read/close per call; benchmark-host side effects drop from ~5861 to ~750 minor faults per read after warmup; isolatednarrow_1000000read_fullwindow 17.8 ms -> 6.7 ms (lab window,exhaustive_cpu_20260807_082144_reader_cache). exhaustive_cpu_20260807_082931_reader_cacherun: fitstableread_fullratio vs astropy 1.365x -> 1.072x;predicate_filteronnarrow_100000016.60 -> 11.01 ms (1.21x -> 1.64x vs astropy); selective 10.68 -> 9.33 ms (1.50x -> 1.71x) via torch-gatherwhere=(mmap off, lab). The_082144/_082931run CSVs were lost in a local workspace reset (noted inbenchmarks.md); the numbers survive in the benchmarks page's generated tables, and the reader-cache direction was spot-verified on current HEAD 2026-08-09 (narrow_1000000read_full: torchfits 7.9 ms vs astropy 16.5 ms, mmap on).- Insert/update rows open the table once per operation (
6d2338c).
Fixed¶
- Read-then-mutate regression: appending rows after a cached read no longer hits CFITSIO error 104 (reader cache eviction is now process-wide, covering cross-thread writers).
- Stale reads: replaced files (new inode) are re-detected via stat identity; cached pread handles are dropped, never serving old data.
- Schema metadata with nanobind >= 2.14 (dropped str->int coercion):
tnull/bscale/bzero/null values parsed from header strings instead of raisingstd::bad_caston the write path. - Tile compression: reject unsupported floating-point payloads on PLIO_1 writes with a clean ValueError.
- Table mutation & schema: strict validation for VLA and object dictionary-table columns.
- Stats transforms accept integer dtypes (BZERO-scaled
read_subsetresults come back as UInt16, which torch cannot reduce): helpers upcast to float32 (int64 keeps precision as float64), masked fills use dtype-safe sentinels, andSigmaClippromotes integer inputs instead of raising.
Dependencies¶
- Vendored CFITSIO at the 1.0.0 tag: 4.6.4 (
extern/VERSIONS.txt). (An earlier revision of this entry claimed 4.7.0; that bump landed after the tag and is recorded under Unreleased below.)
Packaging¶
- Linux wheels now cover x86_64 and aarch64 for CPython 3.10–3.14
(cibuildwheel 4.1; GHA
ubuntu-24.04-armfor aarch64). macOS arm64 is unchanged; runbash scripts/cibuildwheel.shon a Mac to build locally. - PyPI no longer receives an sdist —
pip installcannot fall back to a source compile (the rc5 trap on Ubuntu 26.04 / Python 3.14). - Wheel cmake args no longer point
CUDA_TOOLKIT_ROOT_DIRat conda. CUDA torch at runtime is the same CPU-linked wheel; verify on CANFAR withscripts/verify_wheel_cuda_canfar.sh.
Docs¶
- Benchmark tables refreshed from the 2026-08-07 exhaustive runs (CPU + CUDA);
docs/assets/bench/mirrors the surviving 2026-08-07 CPU/CUDA run CSVs. - Denoise pipeline page with honest dark-vs-bias results and stated limitations; MegaCam CR-cleaning section in the ML guide; examples index row.
1.0.0rc5 — 2026-08-06¶
Fifth release candidate on the 2.13 torch ABI lane; adds prerelease-aware release tooling, a root cache reset entry point, and the macOS compressed-float parity fix.
Added¶
clear_all_caches()— root-level clear of in-process and disk caches (includingcache_root()downloads/samples);clear_cache()stays in-process-only by default.- CI/scripts:
release_lane.py --prerelease rc<N> --applyrenders a lane's release version plus a PEP 440 prerelease suffix (e.g.1.0.0rc5on the 2.13 lane) across all five pinned files;--check/check-laneaccept rc states as the lane base; unit tests cover suffix render/check/reject paths. - Opt-in robust float→int16 packing:
write/write_tensor(..., quantize=)andtable.write(..., quantize=)("robust"or{"lo_q","hi_q","keep_zero"}). Default remains native float (BITPIX=-32/ floatTFORM). - Example:
examples/example_quantize_int16.py. - CLI
compress/decompress: multiple inputs via--out-dir;--split file|hdu(one output per file or per image HDU);-j/--jobs= PyTorch intra-op threads;-J/--file-jobs= multi-file thread pool. - CLI
-J/--file-jobsonverify/stats/arith(and compress). - CLI
arith: image–image operand, multi-HDU stack+ATen, multi-file--out-dir. - CLI Wave 2: batch
copy/transform/cutoutvia--out-dir+-J;statsstd/median;compress --algorithm;header -kwildcards;setkey --delete/@listvia CFITSIOfits_delete_key(keeps compression). - Docs: CPU-only (no CUDA libs) install recipe; “not only for ML” blurb; roadmap 2.0 native engine / GPU-direct (drop CFITSIO).
- Packaging:
torchfits[cpu]/torchfits[cuda]extras (both Linux-only — macOS no-ops, MPS ships in the default wheel) plus one-line CUDA/CPU install recipes (PyPI's default torch already bundles CUDA on Linux x86_64; CUDA builds also run on GPU-less machines via CPU fallback). - Docs: full docs↔code sync review — one-line pinned installs in quickstart /
CLI docs,
write()payload types corrected (no top-level ndarray), CLI-recipe transform kwargs documented, architecture freshness rc5. - CI/scripts:
check-torch-pinsresolves the[cpu]/[cuda]extra pins against the PyTorch indexes on the wheel ABI lane, run as the first step of both CI jobs and the wheel-build workflow so lane drift fails fast (CI lint job, build_wheels tests + wheel jobs,ci-local). macOS passes vacuously (both flavor pins are Linux-only); the doc-drift guard now also requires every extra's exact pin string (e.g.torch==2.10.0+cpu) to appear in install.md / README; unit tests cover the lane guard, the marker-skip path, the missing-extras failure, and the exact-pin doc-drift check. - Lossless compressed float writes: GZIP_1 and integer RICE_1 no longer
silently quantize (
fits_set_quantize_level(0)); float RICE_1 / HCOMPRESS_1 keep CFITSIO default quantization (lossless unsupported), matching astropy/fitsio defaults. Documented inio.write(). - Compression matrix suite (48 tests): all algorithms × int dtypes with astropy oracles, PLIO range rejection, float quantization bound, per-algorithm cutouts, fpack byte-identity.
- Output-parity suite (
tests/test_output_parity.py, 69 tests): bitwise cross-library read parity (torchfits == fitsio == astropy), write round-trips through fitsio for every compression type + quantize + LONGSTRN + uint64 rejection, seeded fuzz-lite sweep. - Bit-faithful write/read fidelity suite (55 tests) with astropy as an independent oracle.
BZIP2_1image codec (vendored CFITSIO patch, opt-in viaTORCHFITS_USE_BZIP2, default ON when libbz2 is available; not part of the FITS standard — astropy refuses it).- CANFAR matrix bench mode: any python × torch lane × device grid from a VOS
wheel bundle (
scripts/launch_canfar_matrix_grid.sh), 41-leg grid run.
Changed¶
open_subset_readermmap path covers unsigned FITS conventions (BZERO/BSCALE).- Warm
read_shapehits shared image-info cache;read_headercaches cards LRU. - Landing one-liner + transparent nav/favicon logos; contributing / release
checklists aligned with verify tiers (
preflight-push/ci-local/release-gate). torchfits headertext mode dumps all HDUs in fitsheader-style blocks.- CLI parallelism docs:
-j(torch) vs-J(file workers). setkey --rename/--deleteremove keywords via CFITSIO delete (no decompressing rewrite);--split hdurejects colliding stems.table.read_torch(..., where=)applies a torch mask after reading projected columns; dialect is simple compare /BETWEEN/ANDonly (full dialect ontable.read).write()no longer rejects numpy arrays: torchfits.write(numpy_array) writes image HDUs (plain, quantized, compressed) instead of raising TypeError. Fixed the public-API bug where docs documented broken behavior.- LONGSTRN/CONTINUE read+write: header values > 68 chars assemble via
CONTINUE chains (
&+ bare CONTINUE) on read andfits_update_key_longstron write. - uint64 image/table writes raise
ValueErrorwith guidance (was an unsupported-column error); the rejection covers numpy table columns too. - TSCAL/TZERO scaling applied to FLOAT/DOUBLE table columns; scaled tables fall back to buffered (non-mmap) reads with physical values.
- int8 images use the BZERO=-128 signed-byte convention instead of raw bytes without BZERO (values ≥ 128 corrupted on read); unsigned table prep registers untouched columns in the synthesized schema.
- ASCII string columns use code-first
Awtforms;fits_schema.parse_tformunderstands the ASCIIAwform soupdate_rowsstops truncating ASCII strings. - LOGICAL decode accepts
'T'/'1'/1(CFITSIO returns converted 1/0 onfits_read_col(TBYTE)); uint64 images (BZERO=2^63) detected as scaled instead of read raw. Header.removefast path for huge HISTORY lists; 20k-delete regression smoke.- TableHDU caches version-gated on the header (were stale after TTYPE/TFORM mutation).
- HCOMPRESS uses 2D 16-row tiles like fpack's default (1D tiles rejected by CFITSIO with status 413).
- Concurrency: Python FITS-cache LRUs and the thread-local metadata cache are lock-protected / size-bounded for concurrent reads; uint16 BZERO offset is fused into the SIMD bswap mmap path.
- Bench docs: multi-host GPU/CPU results from the CANFAR matrix grid.
- Bench docs: rc5 re-run snapshot — CANFAR CPU
exhaustive_cpu_20260806_012620, CUDAexhaustive_cuda_20260806_012651, local CPUexhaustive_cpu_20260806_022603. CUDA 100% fits win rate (smart/specialized), fitstable ≥98.9%; only residual lags are tableread_full/ predicate rows (≤1.15×, fitsio/astropy) and HCOMPRESS_1 (≤1.03×). - Bench labeling: the per-platform results table derives the platform from the benchmark
data (
metadatadevice field /hostcolumn token), not the run-id tag — the local bench script names every runexhaustive_mps_*regardless of platform, so a CPU run on a Linux box used to be mislabeled "macOS arm64 / MPS". The script now tags runsexhaustive_cpu_*on non-Darwin hosts. - SSRF hardening waves: private/loopback/link-local/reserved-address guards
on read_header, read_batch, HDU write paths, and public cpp;
scan_polarsguards before importing optional polars.
Fixed¶
- macOS compressed-float parity: the vendored CFITSIO now builds with
-ffp-contract=off(clang/aarch64 FMA contraction shifted low-ULP bits of decompressed float tiles, e.g. exact0.0read back as ~1.9e-16); the compressed-image parity suite additionally allows ≤1 dtype eps vs fitsio/astropy on macOS, staying bitwise-exact everywhere else. - Table int16 columns with
TSCAL/TZERO: disable CFITSIO auto-scale on read before casting, then apply scale in memory (avoids int16 overflow). - Silent data loss / corruption fixes: compressed float writes silently
quantized by CFITSIO defaults (documented behavior change above); int8
images without BZERO corrupted values ≥ 128; table column ordering scrambled
on rewrite (unordered_map → ordered vectors); ASCII string columns truncated
to one character on
update_rows. - PLIO heap overflow in vendored CFITSIO (
imcomp_calc_max_elemsizing, ASAN- confirmed 16-byte overwrite on incompressible data), fixed by auto-applied patch; bzip2 image codec re-enabled upstream. - LOGICAL (T/F) decode accepts all of
'T'/'1'/1. scan_polarsguarded before importing optional polars.- CI lint packaging dep fix; astropy < 6.0 uint32 tile-compression test failures fixed; py3.10 uint32 astropy oracle skipped below 7.0.
check-torch-pinsmisleading pass on rc states:main()re-initialized thefailedflag after the lane-consistency loop, so a genuine[FAIL]line never gated CI (exit 0), andlane_for_versionrejected therc<N>prerelease suffix — the gate re-fired a spurious FAIL on every rc cut. Suffix is now accepted as the lane base and lane-consistency failures propagate to the exit code (regression tests added).
1.0.0rc4 — 2026-07-20¶
Fourth release candidate for collaborator testing after prep / deep-review cleanup.
Removed¶
- Root aliases
read_table,stream_table,read_table_rows,get_header,get_batch_info— usetable.read_torch/table.scan_torch/read_header/read_batch_info. - Spectral and continuum transforms (
spectral.py,continuum.py) — torchfits keeps FITS I/O–adjacent viz/ML preprocess only; spectroscopy analysis moves out of this package. - Dead
core.py/ChecksumVerifier— checksums go through_Cviachecksum_api. - CLI deprecated aliases
--fitsort,--bytes,--preview— use--keyword-table,--header-bytes,-n/--rows. - Deprecated
table_module=dual-path on cache invalidate/clear. table.to_polars_lazy— usescan_polarsorto_polars(...).lazy().torchfits.cli.rgbshim — importlupton_rgb/write_rgb_imagefromtorchfits.transforms.- No-op
handle_cache=onread_tensorandhandle_cache_capacity=onread_subset(persistent reuse stays onopen_subset_reader).
Added¶
- Skinny metadata:
read_nrows,read_keys,read_shape,read_hdu_type,read_num_hdus,read_colnames,read_extname,read_table_info— CFITSIO structural/key queries without a full header dump. open_table_reader(path, hdu=1)— reusable table handle (mirror ofopen_subset_reader).table.read_torch(..., where=)— fused C++ project+predicate path.FITSHeaderScale.from_path/FITSHeaderNormalize.from_pathvia skinny keys.transforms.as_module/AsModule— thinnn.Moduleadapter fornn.Sequential.- CLI
transform --name Class:key=val,...constructor kwargs. - HTTP Range cutouts + vos/vault remote fetch (prior unreleased work).
- Public
TensorHDU.shape_str/dtype_str; optionalFitsTableDataset(labels=).
Changed¶
read()rejects unknown kwargs withTypeError(no silent swallow of leftovers likepolicy=).- Table
read_torchuses a thin C++ path (skipsread_unifiedimage probes). - Datasets
label_key,get_image_meta, benches/examples lean on skinny meta. - Lazy root
__getattr__uses a lock-backed attribute cache (noglobals()mutation). - Library logger uses
NullHandler(no import-time StreamHandler). - Bench deficit CSV always lists raw lags;
significanceisnoiseorsignificant(floors label only). - Removed disconnected
benchmarks/bench_fast.pyand pixi aliasesbench-fast/bench-fast-stable/bench-core(usebench-fits). - Dead private
read_large_tableleftover and unused_unsigned.py(unsigned paths live in_read_pipeline/write_api/fits_schema). - Table mutation: single
_mutation_cache_barrierpre/post; dtype maps via_ensure_dtype_maps(). FitsTableDataset.__getitem__returns(row_dict, label)formake_loader/fits_collate_fnparity with image datasets.- Lupton examples/docs use
lupton_rgb(r=..., g=..., b=...)(reddest → R). - Example smoke runner auto-discovers
examples/*.py(skips_*.pyhelpers). - DataView BITPIX 64 →
torch.int64; C++write_imagesupports int8/int64.
Fixed¶
- HTTP Range cutouts: NumPy view
byteswap(True)on frombuffer tensor (this torch build has noTensor.byteswap); drop redundant cutout.clone(). - SigmaClip: 0-d
new_zeros(())fill fortorch.where(..., out=)(nozeros_likebuffer). - Filtered table zero-match
where=: keyed empty tensors (not{}). lupton_rgb: Astropy-parity Lupton asinh mapping (per-pixel peak clip). Gallery SDSS / MegaPipe figures regenerated with readable stretch.SubsetReader: uncompressed 2D images mmap the data segment once and slice+bswap into torch (MegaPipe-class mosaics); CFITSIOfits_read_subsetremains the fallback for compressed / scaled / non-2D.- Clearer
TypeErrorwhen inferring FITS TFORM from uint16/uint32/uint64. LogStretch.inverseclamps exponents to avoid float overflow.- Remote prefetch bookkeeping cleaned after completed downloads (locks retained).
- WHERE
BETWEENboundaries no longer use bare\S+(operators/parens). fast_parse_header_cards/Headerkeep empty comments as""(not"None").- Wheel builds pin torch 2.10 ABI (
--no-build-isolation+ before-build install) so release smoke no longer fails against a 2.13-built extension. - Deep-review P0–P4 harden (WHERE OOM gate, batch exception narrowing, HDU close race, prefetch errors, mutation cache barrier, NAXIS overflow guard).
Docs¶
- Removed-names table lists root
read_table/stream_table/read_table_rows/get_header/get_batch_info. - Core I/O cache section documents root vs
torchfits.cachelayers. - Examples:
open_table_reader+ EXTNAMEtable.read_torch. - July 2026 CANFAR/local benchmark refresh: MPS
exhaustive_mps_20260719_143706, CANFAR CPUexhaustive_cpu_20260719_144337, CUDAexhaustive_cuda_20260719_144457; MegaCam20260719_075555; MLml_20260719_145743. - Slim transform gallery; real Lupton RGB figure; removed spectral/continuum docs.
- Core I/O docs point at
table.read_torch/table.scan_torch(root aliases gone). - User Guide ML with FITS: Galaxy Zoo 1 + Legacy Survey one-epoch CNN train; MegaPipe mosaic collage + cutout timing.
- Canonical
TORCHFITS_*env tables in architecture; slimmed duplicates elsewhere. release-gaterunsdocs-contract(example sync + zensical build) anddocs-links(internal hyperlink crawl ofsite/).
Examples¶
example_ml_galaxyzoo_legacy.py,example_megapipe_cutout_collage.py,scripts/fetch_cfht_megapipe_sample.sh.
1.0.0rc3 — 2026-07-18¶
Third release candidate for collaborator testing.
Docs¶
- Readability pass on user-facing pages: rc honesty, corrected migration threading
(private CFITSIO handles since rc2), API notes for EXTNAME / 3D
read_subset, cache vs disk-cache /make_loaderlayering. - Examples gallery: MaNGA LOGCUBE (
example_manga_logcube.py), Lupton RGB from real SDSS g/r/i (example_lupton_rgb_sdss.py, stdlibbz2inflate before read), CFHT MegaCam MEF cutouts (example_megacam_mef_cutouts.py). - Fill API / cache / loader doc gaps:
TORCHFITS_CACHE_DIRvs in-process handle cache, whenoptimize_cacheno-ops on table datasets,make_loadervs plainDataLoader.
Developer workflow¶
- Added
AGENTS.md+JULES.mdagent configuration: weekly automation retargets to bug/perf-only passes; ledger in.cursor/jules-ledger.md; out-of-scope cosmetic PRs are out of scope.
Fixed¶
- String HDU / EXTNAME:
read_tensorandread_subsetaccepthdu="EXTNAME"(e.g.hdu="MYDATA").hdu="auto"still raises a clearValueError. - 3D subset:
read_subset/open_subset_readerpreserve the leading cube axis; window applies to trailing(y, x)only. - Zero-size cutout box: a degenerate width or height keeps the other axis length (no longer collapses both dims to 0).
- Table
where=+ TNULL: filtered reads honorapply_fits_nulls=Trueso sentinel nulls do not leak as real values. - Remote prefetch race:
resolve_local_pathwaits on an in-flight prefetch for the same URL instead of racing a second download onto the same.partial. - Lupton RGB: zero-size bands raise a clear error instead of a cryptic
RuntimeError. .fits.bz2: clearValueErrorwhen CFITSIO cannot read bzip2-compressed paths (decompress first — see Lupton example).
Docs site¶
- GitHub Pages deploys stable (
/, latestv*tag) and edge (/edge/, tip ofmain) fromdocs.yml— use edge to debug docs without a SemVer release. Docs “stable” may be an rc tag; PyPI non-prerelease can lag.
1.0.0rc2 — 2026-07-18¶
Second release candidate on the 1.0 line. CFITSIO concurrent-read correctness,
leftover API/docs/CLI, and cleanup. SemVer 1.0.0 still waits for extended testing.
Install / compatibility¶
- Runtime / build metadata:
torch>=2.10(wheels and pixi stay on the 2.10 ABI lane). Source builds embed the detected torch major.minor as the ABI tag. - Docs: wheels vs source, unified GPU/accelerator install,
configure_for_environmentcalled once at import. Droppedipykernelfrom[dev]. - Disk cache root:
TORCHFITS_CACHE_DIR(default XDG /~/.cache/torchfits); remotes and samples as subdirs; Dataset /make_loaderhonorcache_dir=.
CLI¶
- Short options:
-e/--hdu,-f/--format,-o/--out,-w/--where,-c/--columns,-n/--rows,-k/--keyword(and setkey-k/--key). header --keyword-table(deprecated alias--fitsort).convert --where/--columnsfilter+export; optional FITS table out.probe --header-bytes(alias--bytes) /--timeout.probeSSRF guard: blocks private/loopback/link-local/reserved addresses viagetaddrinfo(all records) and re-validates every HTTP redirect hop.
API / ML¶
read/read_headerdefaulthdu=0(hdu=Nonestill autodetection).- Dataset peers:
FitsTensorDataset(general N-D),FitsImageDataset,FitsCubeDataset,FitsSpectrumDataset(multi-armlayout=, IVAR companions). - HTTP(S) remote prefetch under the configurable cache root.
- mmap guidance for DataLoader / network FS.
- Lift
_HDUInfo/_TableWriteProxyto module scope with__slots__. - HDU
_repr_html_usesscope=col/scope=rowfor accessibility. - Individual HDU HTML reprs: keyboard-focusable container + theme-aware borders (aligned with HDUList/Header).
Correctness / bindings¶
- Concurrent reads open a private
fitsfile*per call (CFITSIO R2); no shared-handle LRU across threads. SharedReadMeta + shared rawfdremain. - Python table/subset paths no longer share one cpp handle across threads.
- Table writes ensure C-contiguous buffers (signed-stride safe) before
fits_write_col; image writes force contiguous host tensors beforefits_write_img. write_table_hduuses RAIIvector<string>forfits_create_tblname/ttype pointers (no mid-throwchar**leak).evaluate_whererejects== NULL/!= NULLon numeric arrays; preferisnull/notnullortable.read(..., where=).- Remove duplicate
num_rowsbinding; clean deadanalyze_tablecomments. append_rows/insert_rowsbest-effort rollback viafits_delete_rowsafter a failed post-insert write.- FITS header integer keys use
PyLong_Check+ overflow-checkedTLONGLONG. - Empty-primary MEF compressed write HDU indexing fixed.
- Automation integrations: probe SSRF (#216), hoist inner classes (#214), HDU HTML a11y (#213/#219).
Benchmarks / tests¶
- ML loader: on-disk
size_mbfor compressed cases; pin compressed torchfits tohdu=1(matches fitsioext=1); missing numeric CSV fields useNone. - GPU transports: pass
quick=into table_build_cases. - Table filter tests assert exact fixture row counts.
- Concurrent same-file image/table read smoke tests.
- Benchmarks: new local MPS run
exhaustive_mps_20260718_180230; Linux CPU/CUDA hosts remain the rc1 benchmark runs until CANFAR is re-driven against this tag.
Docs / transforms¶
- Mermaid diagrams in architecture (zensical superfences).
- Architecture: per-read handles, deliberate skip of CFITSIO iterator/
where. - Roadmap: CFITSIO 1.1 leftovers + permanent design choices from the design review.
- Advanced transforms frozen for 1.0; Lupton RGB wrapper in
transforms.lupton_rgb; richer multi-band RGB deferred to 1.1. - MegaCam cutout bench: ZNAXIS-aware HDU discovery; peer fitsio ranking; materialize-once baseline; payload-based throughput (not whole-file MB/s).
- Docs landing: lean browse grid; nav uses mark-only logo (
torchfits-logo-mark.png). - Vendored CFITSIO docs pin 4.6.4; note
fits_iterate_dataintentionally unused. - cfitsio-direct: Rice + optional MegaCam
cutout_repjobs.
1.0.0rc1 — 2026-07-17¶
Release candidate for the 1.0 API. SemVer 1.0.0 waits for post-rc extended testing; do
not treat this tag as the final 1.0.0 freeze.
Changed¶
-
verify/verify_checksums: missing checksum keywords are success. Files withoutDATASUM/CHECKSUMnow returnok=True,status="no_checksums", CLI textOK (no checksum keywords), exit 0. Previously CLI exited 4 (FAIL). Aligns withfitsverify(missing keywords are not corruption). Scripts that treated any nonzero verify exit as “bad file” must key offstatus == "fail"/ exit 4 instead. -
Root table helpers deprecated.
read_table,stream_table, andread_table_rowsemitDeprecationWarning. Prefertorchfits.table.read/read_torch/scan_torch.
Fixed¶
-
ArcsinhStretch/LogStretch: validatea > 0in__init__. Previouslya=0silently producedNaN(div-by-zero ininverse/forward). Now raisesValueErrorwith a clear message. (transforms/stretch.py) -
_normalize_row_slice: reject negativestopwithValueError. Previouslyslice(0, -1)silently returned 0 rows — the function cannot resolve negative indices without knowing the total row count. Now raisesValueErrorwith an actionable message. (_table/utils.py) -
Empty
WHERE/row_slice/rowsresults: preserve column schema. Previously all empty-result paths returnedpa.table({}), losing all column names/types and causingKeyErroron valid queries (e.g.where="ID > 9999"on a table with no matches). New_empty_table_with_schema()helper builds typed empty tables from FITS header cards, preserving requested column ordering. When header schema is unavailable but columns were requested, returns null-typed empty columns instead of{}. (_table/read.py) -
io.write()header type: widen toHeader | dict[str, Any] | None. RemovedTODO(1.0)andtype:ignore[arg-type]in_table/write.py. The runtime already accepted dicts; only the type annotation was narrow. (io.py,_io_engine/write_api.py,_table/write.py) -
Example runner:
REQUIREDexamples can no longer silently skip. OnlyOPTIONALexamples (e.g.example_polars.py) may skip on missing deps.REQUIREDexamples always surface failures. (examples/test_examples.py)
Added¶
docs/cli.md:### verifysection (three labels, exit codes, fitsverify note); CLI cold-start / process-tax note.docs/compatibility.md: Python / PyTorch / Arrow / platform matrix.scripts/clean_install_smoke.sh: local wheel → fresh venv install smoke.tests/test_http_probe_fixture.py: Range HTTP replay forprobe.- HTTP probe JSON records include
"source": "http"(matchesvosprobe). - Tests: stretch
a<=0, empty schema preservation, verify messaging contract, deprecation warnings, HTTP probe fixture.
Benchmark evidence¶
- Multi-host benchmark results (from b1 same-day refresh, still current for rc1):
exhaustive_mps_20260717_040150,exhaustive_cpu_20260717_040146,exhaustive_cuda_20260717_042840. - Local release-suite (
20260717_212321, Mac MPS, mmap matrix,--no-gpu): 2,825 rows, 3 deficit rows, exit 0. No domain failures.
Validation¶
848+ tests; mypy / ruff clean; docs integrity; examples runner REQUIRED green;
bash scripts/clean_install_smoke.sh; HTTP probe fixture.
1.0b1 — 2026-07-17¶
Beta freeze of the public FITS → tensor / dataframe story. Not a SemVer 1.0.0 API freeze (rc line followed for extended testing + blockers).
Added¶
torchfits.table.read_torch(tensor-column dataframe path) andtable.read_arrow(synonym oftable.read).- Docs gallery: KaTeX math, transform before/after figures, CLI recipes, real-sample cache helpers (merged via docs gallery work).
- Release reviews: rendered docs, API adoption, deep code, real-data CLI vs astropy/fitsio/gnuastro/CFITSIO (FITSH skipped).
Changed¶
- Docs teach FITS tables as dataframes while keeping the
torchfits.tablenamespace; which-reader box demotes compatibility aliases. - Landing / site_description: tensors and dataframes (columnar catalogs).
Fixed¶
torchfits transformon integer HDUs: promote to float before transform and write float outputs without reusing integer BITPIX headers.
[0.9.3] — 2026-07-17¶
Added¶
torchfits header --fitsort --keyword …multi-file keyword table (same idea as qfitsdfits | fitsort).- Optional
vos:/vos://probe when thevospackage is installed. - Invalid
--hduvalues exit with usage code 2 instead of a traceback. - Lean
_repr_html_onTensorHDU,TableHDU, andTableHDUReffor notebooks. torchfits convert --to pngLupton RGB preview via stdlib PNG (no Pillow / NumPy). PPM removed.- Table convert formats: parquet, csv, tsv, and arrow (Arrow IPC / Feather V2). Streaming writers for large catalogs (CSV/TSV: flat columns only).
Changed¶
torchfits.transformsis a package split by domain (stretch,normalize,fits_meta,spectral,continuum,clip) with the same public__all__.- Transforms docs: not
nn.Module; instance-local inverse state; Advanced notes forBandMath,PhaseFold,AsymmetricLeastSquares,AlphaShapeContinuum; invertibility + helpers tables. - Parquet convert uses streaming
write_parquet(..., stream=True)(out-of-core). - Multi-host benchmark refresh (
exhaustive_mps_20260717_040150,exhaustive_cpu_20260717_040146,exhaustive_cuda_20260717_042840): CUDA 0 deficits, CPU 1, MPS 16. scripts/gpu-bootstrap.shpinstorch>=2.10,<2.11so CANFAR cu128 installs do not pull PyTorch 2.11 and fail the ABI gate.
Fixed¶
- Block CFITSIO
sh://filenames (command injection via/bin/sh), extending the existing|checks.
[0.9.2] — 2026-07-16¶
Added¶
torchfitsCLI — MEF-aware shell tools:info,header,verify,diff,stats,table,convert,copy,arith,cutout,compress,decompress,transform,probe,setkey. JSON/JSONL output and stable exit codes. Guide:docs/cli.md.
Changed¶
- Public imports — root is I/O + HDU only. Import transforms from
torchfits.transforms.torchfits.hduis a documented namespace. - Removed
read_fastandread_image(useread/read_tensor). Deleted the unused_fastiomodule. - Table policy helpers (
can_use_*, …) are no longer listed intable.__all__.
Fixed¶
- Signed-byte (
BZERO=-128) and unsigned smart device reads convert on the host then copy once to CUDA/MPS. read_subset/SubsetReaderkeep signed-byte and unsigned integer conventions as narrow dtypes (int8/uint16/uint32) instead of float-promoting every cutout.- Automatic table
where=withmmap=Trueuses native mmap-scan pushdown when safe;mmap=Falsereads then filters in Arrow/tensor space. - CFITSIO
MINDIRECTreset to 8640 so ~13 KB HCOMPRESS tiles use direct tile I/O. - Multi-byte mmap image reads use NEON/SSSE3 endian convert for all sizes.
- Uncompressed BYTE_IMG reads use direct
preadfor mmap on and off. - One-shot image reads use thin
cpp.read_fullinstead of handle-cache scaffolding on the cold path. - Repeated cutout benches use the persistent subset reader (open once).
- Deficit table: images any lag above ε; Arrow tables allow ≤1.05×; fitsio
excluded from mmap-on peers. Linux CPU/CUDA strict-gate 0 deficits; Mac
MPS 4 on
exhaustive_mps_20260717_000853.
Docs¶
- Site logo/favicon:
torchfits-logo.png. - README / benchmark run IDs aligned with
docs/benchmarks.md.
[0.9.1] - 2026-07-14¶
Fixed¶
- Native wheel metadata now constrains PyTorch to the 2.10 ABI used to build the extension. Torchfits 0.9.0 incorrectly allowed newer incompatible libtorch releases, which could segfault during image or table conversion.
- Native builds and imports now reject mismatched PyTorch ABIs, and every CI build path installs the same PyTorch minor used by the release wheels.
0.9.0 - 2026-07-14¶
Fixed¶
- Writing one FITS file no longer invalidates borrowed native handles for unrelated files. Native cache clearing now defers closing in-use handles, preventing a subsequent read from dereferencing a closed CFITSIO handle.
- Atomic table-column rewrites now close every managed
HDUListborrower for the target path, so nested open contexts cannot retain an old inode and erase an earlier mutation. write()normalizesos.PathLiketargets before native cache invalidation.- Wheels no longer include the C++ build-source directory.
- Numeric tensor-to-Arrow conversion now shares the tensor's NumPy buffer instead of iterating through PyTorch storage one byte at a time.
- Automatic table predicates use the fast native full-read path followed by
Arrow filtering; native row-wise pushdown remains available through the
explicit
backend="cpp"policy.
Added¶
read_polars()— one-call FITS-to-Polars convenience function. Callsread()withinclude_fits_metadata=True, converts viapl.from_arrow(rechunk=False), and returns aFITSPolarsFramewrapper that preserves FITS column metadata (TFORM, TUNIT, TDIM, TNULL, TSCAL, TZERO) alongside thepl.DataFrame. Delegates__getattr__,__getitem__,__len__to the wrapped DataFrame.scan_polars()— genuine streaming Polars path. Yieldspl.DataFramebatches viapl.from_arrow(batch, rechunk=False)overscan(), without materializing the entire Arrow table. Unliketo_polars_lazy(), no full table is built.FITSPolarsFrame— lightweight dataclass wrapper aroundpl.DataFramewithfield_metaandtable_metadicts for FITS metadata preservation.- Transform masks now thread through FITS-aware normalization and clipping; spectral resampling uses torch-native interpolation with parity references for vectorized continuum, phase-folding, wavelet, and sigma-clipping paths.
- CANFAR CUDA exhaustive (
exhaustive_cuda_0.9.0_20260714_065950) — 3,648 normalized rows across the mmap on/off and CUDA matrix; 7 deficits, all at or below 1.439×, with no large-N deficit.
Changed¶
- Removed the never-implemented
TensorHDU.stats()and its empty native result from the supportedtorchfits.cppinventory instead of inventing statistics semantics during the 0.8 API freeze. ci-localnow runs its pre-build package-isolation checks againstsrc/, so a clean Linux clone no longer depends on a pre-existing editable install.- Native cache environment limits are validated before loading Torch or the extension module, preserving useful configuration errors in clean installs.
- Scoped extension-only visibility and semantic-interposition optimizations to
_C; applying them directory-wide also changed vendored CFITSIO's C ABI and aborted Linux ASCII-table writes. - Raw, unmapped image reads now support FITS
BITPIX=64images astorch.int64, matching the mapped and scaled readers. - GitHub workflows use the Node 24-based
actions/checkout@v5andactions/setup-python@v6, and pin Apple Silicon testing tomacos-15instead of following the rollingmacos-latestmigration. -
Removed the environment-dependent optional
torch_frameinheritance fromTableHDUand thetorchfits.hdu.TensorFramealias. FITS table columns stay as tensor/list mappings, Arrow is the interchange boundary, and Polars is the dataframe surface. Any legacy dataframe bridge remains outside torchfits. -
rechunk=Falsedefault onto_polars(),to_polars_lazy(),scan_polars(),read_polars(), and top-levelto_polars(). Avoids Polars' unnecessary chunk concatenation when Arrow data is already single-chunk (the common case fromread()). Passrechunk=Trueexplicitly to restore the old behavior. to_polars_lazy()docstring — clarified that it materializes the entire Arrow table eagerly before wrapping asLazyFrame. Users seeking true streaming should usescan_polars()instead.
Removed¶
"cpp_numpy"table backend alias — the deprecation alias introduced in 0.7.0 is removed. Passbackend="cpp"instead of"cpp_numpy". TheDeprecationWarningis now a hardValueError.should_skip_cpp_numpy_for_where— internal alias removed fromtorchfits._table_engine. Useshould_skip_cpp_for_where.
0.7.0 - 2026-07-11¶
Added¶
FitsTableIterableDataset— constant-memory table streaming viatable.scanwith worker sharding by scan batch index.FitsCutoutDataset— map-style patch training from(path, hdu, x, y, …)cutout specs.- Zensical documentation site —
zensical.toml,docs/index.md, GitHub Pages workflow, andpixi run docs-build/docs-serve. migration_datasets.md— breaking-change guide for removed legacy datasets.transforms.__all__— explicit public transform catalog.- CI
release-gatejob — upstream parity, docs contract, data, transforms, and security smokes on Python 3.13. - Lab benchmark refresh (
exhaustive_0.7.0_20260711_022156) — full exhaustive lab run (3516 rows, mmap matrix + MPS); CPU performance floor unchanged (core deficits ≤1.33×). - CANFAR CUDA exhaustive (
exhaustive_cuda_0.7.0_20260711_055635) — 3626 rows, 11 deficits on staging GPU; artifacts archived tovos:sfabbro/torchfits-gpu-bench/. - CANFAR bench launcher — headless GPU sessions on staging with VOS persistence
via
vcp(scripts/launch_canfar_gpu_bench.sh,scripts/fetch_canfar_bench_vos.sh).
Changed¶
- Torch-first
table.readC++ path —backend="cpp"reads viaread_fits_table_rows/TableReader.read_rows(torch tensors) instead of the numpy hop; Arrow conversion stays at the PyArrow boundary only. - Table backend rename — public backend
"cpp_numpy"renamed to"cpp"; the old name still accepted withDeprecationWarning. - Legacy datasets removed —
torchfits.FITSDatasetandtorchfits.IterableFITSDatasetdeleted; usetorchfits.datatyped datasets. table.pytrim — re-exports public API only (private_helpers no longer re-exported fromtorchfits.table).- Package description — PyPI/README positioning for ML datasets + transforms.
Removed¶
src/torchfits/datasets.py— superseded bytorchfits.data.
0.6.0 - 2026-07-09¶
Changed¶
- Unified C++ table chunk reads: Refactored
_read_cpp_numpy_tableto clean up the 7-deep C++ dispatch fallback chain andhasattrchecks, delegating directly to the modern C++TableReaderandread_fits_table_rows_numpyAPIs. This successfully resolves Roadmap Track B1. - Version synchronization: Unified package version triplet to
0.6.0acrosspyproject.toml,pixi.toml, and package source. - Blocking mypy in CI — the
mypy src/step in GitHub Actions is now a hard gate (previously non-blocking via|| echo). All 103 type errors have been resolved across 18+ source files. Added[[tool.mypy.overrides]]inpyproject.tomlforpyarrow.compute(attr-defined) andpyarrow.*(ignore_missing_imports).
[0.6.0b2] - 2026-07-09¶
Added¶
- Predicate filter improvements (all sizes now use C++ pushdown): The
predicate_filterpath delegates to C++ for all table sizes, eliminating the Python fallback for narrow tables. Theread_policy.pysize threshold is removed — safe non-VLA tables always use C++ pushdown. Narrow-table predicate_filter lag vs fitsio reduced from ~2.86× (smallest) to ≤1.07×. - Lightweight is_compressed check: Compressed images use a fast O(1) header probe instead of opening and parsing the full HDU, reducing overhead on batched compressed-image reads.
- Thread-safe caches: CacheManager internal data structures use
std::shared_mutexfor concurrent reader access, safe under multi-worker DataLoader patterns without global GIL serialisation. - Parallel scan with sequential fallback: The C++ mmap pushdown scan is now
parallelised via
at::parallel_forwhentorch::get_num_threads() > 1, with a zero-overhead sequential path when single-threaded. Addedposix_madvise(POSIX_MADV_SEQUENTIAL)to the filtered scan path for kernel prefetch hints. - Lab benchmark refresh (mmap-on+off, 0.6.0b2): 2754 rows, 3 deficits
in
20260709_163739— down from 0.6.0b1's 14 deficits and 0.5.0b4's 22 deficits. Remaining 3-deficit breakdown: - 3 fitstable (narrow):
predicate_filteronnarrow_{10000,100000,1000000}(1.07–1.25× behind fitsio;narrow_1000dropped below the deficit threshold). The gap is now dominated by Python dispatch + Arrow conversion overhead, not the C++ scan itself (which reaches near-parity with fitsio at ~11.4 ms vs ~11.0 ms for 1 M rows). All compressed-image deficits eliminated. The uint16/uint32 mmap-on regression that motivated 0.5.0b4's bswap+BZERO merge is no longer in the deficit table. - Multi-worker DataLoader coverage:
tests/test_data.pynow exercisesmake_loader(..., num_workers=2)for bothFitsImageDatasetandFitsImageIterableDataset. Tests fork a subprocess to keep CFITSIO's threadpool away from pytest's own threadpool, and verify that every file is seen exactly once regardless ofnum_workersand shuffle seed. - End-to-end FITS round-trip coverage in
tests/test_transforms_e2e.py: TestEndToEndImageRoundTrip— write / read / scale / inverse for INT16 with custom BSCALE/BZERO, BZERO=32768 unsigned convention, and INT32 with rescaling.TestEndToEndTableRoundTrip— FITS binary tables with TSCAL/TZERO use real on-disk encoding (via astropy), thenFITSScaleColumns.from_header+TNullToNan.from_headerround-trip is verified to within the storage precision.TestEndToEndFITSHeaderNormalize— full int16 BZERO=32768 round-trip through the header-driven normaliser.- Release-gate now includes
tests/test_data.py,tests/test_transforms.py, andtests/test_transforms_e2e.py. This closes the torchfits.data documented with multi-worker test coverage and torchfits.transforms round-trip tests for scaled images and tables gate items. AsymmetricLeastSquares(lam, p, max_iter, dim)— Eilers 2003 penalised baseline correction with asymmetric weights. Iteratively solves the Whittaker smoother(W + λD^T D)z = Wywith differential weighting (p above baseline, 1-p below). Standard in Raman/NIR spectroscopy. D^T D penalty matrix built in float64 for numerical stability at large λ. Additive decomposition (invertible).AlphaShapeContinuum(half_window, iterations, dim)— Morphological closing (dilation→erosion) viaunfold+ max/min. Produces a guaranteed upper envelope (always ≥ signal). Practical approximation to the full alpha-shape algorithm (RASSINE). Additive decomposition (invertible).AsymmetricSigmaClip(n_low, n_high, dim)— Simple one-pass asymmetric sigma-clipping outlier rejection usingestimate_background(median + MAD). Supports different lower/upper sigma thresholds; replaces outliers with per-group median. Lossy (no inverse)._build_d2_matrixinternal helper for the n×n pentadiagonal second-difference penalty matrix D^T D used by the Whittaker smoother / AsLS.- 27 new tests for the three transforms (201 transforms tests total, all passing).
- Example coverage for
AsymmetricLeastSquares,AlphaShapeContinuum, andAsymmetricSigmaClip(later removed with the spectral/continuum hard-cut). - All three transforms exported to the root package for direct
from torchfits import AsymmetricLeastSquaresaccess. - Documentation for all three transforms in
docs/api.mdandREADME.mdtransform tables.
0.6.0b1 - 2026-07-08¶
Removed¶
- Removed deprecated
read_large_tablefunction (usestream_tableorread_tableinstead). - Custom WHERE AST evaluator runtime (~120 lines) from
_where.py:_evaluate_cmp,_evaluate_in,_evaluate_between,_evaluate_isnull,_evaluate_where, and theevaluate_wherepublic alias. Replaced withpyarrow.computenative predicates via_where_mask_for_table. The parser, tokenizer, normalizers, andwhere_columns_from_aststay for C++ pushdown path compatibility. - Compressed parallel decompression path (~350 lines):
try_read_compressed_rows_parallel,compressed_parallel_enabled/min_pixels/min_rows_per_thread/max_threads/hcompress_enabledhelpers,load_bswaptemplates,FitsHandleGuardlocal class,is_parallel_compressed_codec_cached,compressed_parallel_cache, andhardware_concurrencydependency. CFITSIO's built-in decompression already covers this serially — the 2-thread cap meant the heuristic rarely activated. - Unused
read_rice_parallel(~320 lines) fromcompression.cpp— vendored Rice decompression, nanobind binding, and the entirecompression.cpp/compression.hfiles. Dead after compressed parallel path removal. bind_compressionfrombindings.cpp— only boundread_rice_parallel.
Changed¶
- 3→1 C++ read path merge: Extracted a single
read_tensor_canonical()infits_detail.hand converted three read paths (read_full_cached,read_full_nocache,FITSFile::read_tensor) into thin wrappers, eliminating ~455 lines of duplication. - API naming consistency: Renamed
read_image_canonical→read_tensor_canonicalandFITSFile::read_image→FITSFile::read_tensor, aligning C++ with the Pythonread_tensor/write_tensorAPI. - bswap+BZERO merge: Merged the two-pass byte-swap and BZERO offset into a single
parallel_forin the multi-byte mmap fast path. For unsigned images (uint16 with BZERO=32768, uint32 with BZERO=2147483648),bswap + addexecutes in one traversal instead of two. - Unsigned mmap fast path unlocked:
_read_unsigned_image_if_needednow defers to the C++ path whenmmap=True, lettingread_tensor_canonicalhandle unsigned conventions natively (single-pass bswap+BZERO returning uint16/uint32 directly). Previously Python preempted C++ by callingread_full_rawand doing a second offset pass — making the bswap+BZERO merge dead code. uint32_2d: 8.3× faster (now beats fitsio); uint16_2d: 3.5× faster; 5 deficits eliminated. - Vectorized string decode: Replaced per-row Python
forloops ininterop.py(to_pandas,to_arrow),table_hdu.py(get_string_column,to_fits), andtable_hdu_ref.py(get_string_column) withnp.char.decode()+np.char.rstrip()for significant speedup on large string columns. - Deduplicated
fits_schema.py:column_tnull_map()delegates to_iter_tfields_indexed()instead of reimplementing the TTYPE/TNULL iteration loop. - Deduplicated unsigned dtype and TFORM parsing:
_table/read.pynow delegates tofits_schema.unsigned_column_dtypes_from_header()andfits_schema.iter_table_columns()instead of reimplementing TZERO/unsigned detection and TTYPE/TFORM header walks. - Table schema fast path:
table.schema()skips data reads whenwhere=None, inferring the Arrow schema directly from FITS TFORM header cards (≤1 header pass). - C++ source extraction: Split
fits.cpp(4552 lines) intofits_detail.h,fits_file.h/.cpp,fits_rw.h; splittable.cpp(3432 lines) intotable_types.h,table_reader.h(header-only),table_mutation.h/.cpp. Removedextern "C"linkage from table mutation functions to fix UB from C++ exceptions crossing C ABI boundaries. Removed dead declarations, unused types, stale comments, and double includes. - Merged cache stats:
CacheManager.get_stats()now pulls I/O engine metrics (io_hits,io_misses,io_total_requests) from the cache subsystem. - WHERE evaluator → Arrow compute:
TableHDU.filter()now builds a minimal Arrow table and delegates to_where_mask_for_table(pyarrow.compute native predicates) instead of running the old NumPy-based custom evaluator. The parser stays for C++ pushdown path compatibility. - Table read unification: Extracted
_read_ranges_as_chunkfrom_read_cpp_numpy_tableinto shared_table/engine.py, removing ~50 lines of duplicated code. - CI: Added non-blocking
mypy src/step to the GitHub Actions lint job. - Benchmark fairness fix: fitsio is no longer unconditionally skipped — runs when
mmap=offfor fair buffered-read comparisons (449 fitsio OK rows in fits domain, 180 in fitstable). examples/example_image_dataset.py:optimize_for_dataset+ correctpin_memorywhen reading directly to CUDA.scripts/run_exhaustive_bench_and_patch_docs.shskips rebuild when extension imports.
Fixed¶
- Root I/O attributes now resolve to the actual public functions, preserving
inspectable signatures, tracebacks, and identity while keeping bare
import torchfitsfree of PyTorch, NumPy, Arrow, and the native extension. torchfits.cppnow has an explicit FITS-native__all__; future compiled symbols no longer become public accidentally. Direct attribute delegation is retained for pre-1.0 compatibility.- Every lazy root export now has a matching
TYPE_CHECKINGdeclaration, so the shippedpy.typedmarker covers the complete documented root API. - Removed the empty, misleading
cacheextra: adaptive cache sizing uses the standard library. Documentation now states that PyArrow is the core table runtime while Pandas, Polars, and DuckDB are optional. - Runtime initialization no longer swallows native-load or invalid cache configuration errors and then marks the failed initialization as complete.
- Header-card write failures and HDU header-preservation failures are no longer silently ignored; callers now receive the native error instead of a successful return with lost metadata. A dead duplicate header helper was removed.
- Overwriting an existing FITS file is now transactional: the complete replacement is written beside the target and atomically installed only after success. Validation or native-write failures preserve the original bytes and file mode instead of deleting the user's file.
- HDU insert, replace, and delete operations use the same transactional rewrite rule, so a partial multi-HDU rewrite cannot replace the original file.
- Iterable HDU writes reject empty sequences, unsupported objects, header-only dictionaries, and non-tensor image payloads instead of silently emitting empty HDUs.
- Image datasets now route through the unified image reader, so their documented
mmap="auto"policy works instead of reaching the bool-onlyread_tensorboundary. Remaining immutable column tuples are normalized at public list boundaries. TableHDUvalidates its trust boundary: non-mapping inputs and columns with inconsistent row counts fail immediately instead of creating an internally inconsistent table.TableHDU.from_fits()now uses the publicread_table()pipeline instead of opening a separate native table/header path, keeping cache, validation, and runtime initialization behavior consistent with the rest of the package.- Removed the duplicate
whereentry from the package root__all__contract. - Scoped mypy's missing-import exceptions to optional dataframe integrations
and the compiled extension, allowing real Python type errors to
surface. Mypy now checks untyped function bodies and is a blocking local
preflight and CI check; the resulting
TableHDURefcolumn-sequence mismatch was fixed at the Arrow boundary. - Release wheels now run image and table round-trip tests against the installed artifact, and macOS arm64 wheels use the platform's real minimum deployment target (11.0). Platform documentation now matches the wheel matrix. Vendored CFITSIO remains statically linked but its development headers, archive, and CMake/pkg-config metadata are no longer copied into wheels.
- Multi-worker DataLoader tests now use a real
__main__guard, matching the macOSspawncontract instead of recursively creating workers frompython -c; timeout failures preserve worker stderr for diagnosis. - Vectorized NULL evaluation:
_where.pyreplaced Python-loopnp.array([v is None for v in val])with vectorized(val == None)for element-wise null checks. - Fixed unused imports in
tests/test_cache_config.py. - Security: Block CFITSIO pipe injection bypass via leading
!prefix (!|command) incheck_fits_filename_security; also enforced on unified cache open path. - GPU
scale_on_devicepreserves narrow integer H2D for FITS signed-byte (int8) and unsigned uint16/uint32 conventions instead of promoting through float32 or int64 on CPU.
Added¶
- Header: O(N) construction for large dict inputs via keyed fast-path in
_set_card(2000 keys ~0.002s locally vs ~2.5s pre-fix). - Jupyter: Scrollable, sticky-header HTML repr for
HeaderandHDUList. tests/test_scale_on_device.py— signed-byte, unsigned, and fitsio parity checks.- Release gate includes
test_scale_on_device.py. .cursor/skills/release-api-freeze-review/— pre-tag API/feature freeze review workflow._table/engine.py— shared C++ table read dispatch module with extracted_read_ranges_as_chunkhelper (de-duplicated from_read_cpp_numpy_table).
Performance notes¶
- Local
bench_ml_loader.pydiagnostic (30×512² float32, CPU, 2 epochs): Rice-compressed 1.12× vs fitsio; uncompressed within ~4% (tune handle cache for your file count). - Lab exhaustive refresh (
exhaustive_mmap_0.5.0b4_20260630_162835, H100 MIG): 3626 rows, 13 deficits (down from 22). Integer CUDA gaps closed; remaining are marginal int8 (≤1.2×) and coldlarge_uint32_2dCPU vs astropy (~1.5×). - User-profile refresh (
unsigned_mmap_fix_20260708, CPU, mmap=on): 1,377 rows, 25 deficits.torchfits_specializeduint32_2d now beats fitsio (was 5–10× behind); uint16_2d at ~1.6× vs fitsio (was ~5×). Remaining deficits dominated by medium-size unsigned reads and compressed HCOMPRESS. Torchfits dominates table I/O (886×–2,318× vs astropy), image reads (7.92× vs astropy, 1.76× vs fitsio on large float32), and repeated cutouts (17× vs astropy, 1.09× vs fitsio).
[0.5.0b4] - 2026-06-30¶
Changed¶
- Centralized FITS binary-table header parsing in
fits_schema(TFORM/VLA/string/bit/unsigned). table.readno longer recurses forwhere=; strategy lives in_table_engine.read_policy.- Table C++ handle caches moved to
_table.cache; I/O cache invalidation no longer depends on importingtorchfits.table. - README highlights 0.5.0 features and published benchmark speedups; API docs document table
backends and
where=tuning environment variables.
Added¶
- Unit tests for
fits_schema, table where-read policy, and runnable example scripts. - Public
torchfits.table.TABLE_BACKENDSconstant. pixi run release-gatetask matching the release checklist parity/docs/examples gates.
0.5.0b3 - 2026-06-30¶
Changed¶
- Refocused torchfits as a FITS I/O package: images, HDUs, headers, checksums, compression, FITS tables, caching, and table interop.
- Removed stale public claims that torchfits owns WCS, sphere geometry, HEALPix, sky-domain simulation, or training pipelines. Those domains belong outside torchfits.
- Added a roadmap and compatibility matrix that distinguish supported, partial, unsupported, and out-of-scope behavior.
- Replaced broad parity claims with test-backed parity tiers for common fitsio, Astropy, and selected CFITSIO-backed workflows.
Added¶
- Extended benchmark matrix: native uint16/uint32 2D image fixtures, typed binary tables (BIT/complex/string columns), and ASCII table fixtures.
bench_all.py --mmap-matrixruns mmap-on and mmap-off passes in one CSV so the I/O transport table can populate bothdisk→CPUanddisk→RAM→CPU(plus GPUdisk→CPU→GPU/disk→RAM→GPUwhen CUDA/MPS is available).scripts/run_exhaustive_bench_and_patch_docs.shfor lab-profilebench-allon CUDA/MPS hardware with automaticdocs/benchmarks.mdrefresh.- Lab CUDA benchmark snapshot
exhaustive_mmap_0.5.0b3_20260630_063118(3474 rows, mmap on+off matrix, 720 GPU transport rows on H100). docs/parity.mdfor the public compatibility matrix.- Astropy upstream smoke coverage for common image, HDU, compressed-image, table, ASCII table, VLA, complex column, and scaled-image workflows.
- Documentation integrity checks for stale WCS/sphere/HEALPix ownership claims.
- Supported-status promotion for in-place mmap table updates on COMPLEX
(
1C/1M), BIT (8X), and fixed-width STRING (12A-style) columns.torchfits.table.update_rows(..., mmap=True)now writes these column types correctly on disk. Verified via raw byte inspection and an astropy upstream-reader roundtrip. VLA columns remain explicitly unsupported in the mmap fast path by design. - Astropy and fitsio upstream smoke coverage that exercises the
COMPLEX / BIT / fixed-width STRING mmap-update parity shift,
including right-padding to the declared column width and verification
vs the upstream readers. The 8A-string assertion falls back to
astropy because the local fitsio upstream misdecodes updated
8Arows (the on-disk bytes are bit-exact to the expected layout; this is an upstream-reader limitation, not a torchfits writer bug). tests/test_astropy_upstream_smoke.py::test_astropy_compimage_compression_variants_match_torchfitsexercising additionalastropy.io.fits.CompImageHDUcompression variants (RICE / HCOMPRESS / PLIO) round-tripped against torchfits.
Fixed¶
- API docs and install guide now reference
torchfits.cachefor cache tuning (configure_for_environment,get_cache_stats,clear_cache) and the root I/O helpersget_cache_performance/clear_file_cachewhere appropriate. - Roadmap mmap limitations updated to match the parity matrix (BIT and fixed-width STRING mmap updates are supported; VLA and scaled columns remain partial).
Removed¶
- Dataset/training helper namespace from the torchfits package contract.
0.5.0b2 - 2026-06-30¶
Fixed¶
- Patched
fitstablespecialised column projection and row slicing benchmark errors due to invalidpolicyargument. - Cleaned up C++ build flags in
bench-gputo remove strict CUDA and Torch pins. - Reviewed C++ codebase for potential memory leaks, redundant hardware heuristics, and API bounds.
Added¶
- Restored core FITS benchmarks from v0.3.2: ML DataLoader performance (
bench_ml_loader.py) and GPU Memory usage/leak validator (bench_gpu_memory.py). - Added exhaustive progress print logging during benchmark execution.
- Added persistent cutout / multi-cutout repeated read benchmarks (
SubsetReader/open_subset_reader) for both CPU and GPU. - Added
read_tensorfor reading N-dimensional arrays (1D spectra, 2D images, 3D cubes, xD arrays) directly to a single PyTorchTensor. - Added
write_tensoras the specialized PyTorch-native writer for writing single PyTorchTensors directly to FITS files.
Deprecated¶
- Deprecated
read_imagein favor of the more general and PyTorch-nativeread_tensor.
0.5.0b1 - 2026-06-29¶
Changed¶
- Repository home:
github.com/astroai/torchfits. - Default development Python is 3.13 (pixi); supported install range remains 3.10+.
- Development Status classifier promoted to Beta.
- Removed obsolete diagnostic benchmarks, scratch scripts, and legacy HEALPix/WCS artifacts.
- CI rewritten: ruff-only lint, multi-OS/Python test matrix, CFITSIO vendoring via
extern/VERSIONS.txt. - Wheel builds: portable flags (no
-march=native),cp310–cp313on macOS and Linux.
Added¶
- GPU I/O transport benchmark rows (
bench_gpu_transports.py) with MPS on Apple Silicon and CUDA on Linux. pixi run bench-mpsfor Apple Silicon accelerator benchmarks.- Automated benchmark report workflow (
.github/workflows/bench-report.yml). scripts/render_bench_deficits.pyfor documenting performance deficits without fixing them.
Fixed¶
- Table mutations now invalidate FITS path caches via internal
iohelper (fixestorchfits._invalidate_path_cachesAttributeError).
Earlier releases¶
Earlier 0.1.x through 0.3.x releases included broader experimental astronomy domains. The current package contract is FITS I/O only; consult the current README, API reference, roadmap, and parity matrix for supported behavior.