Project Roadmap¶
The vision, planned milestones, and architectural evolution of torchfits.
Released¶
1.0 — Foundation (2026-08-09)¶
The 1.0 release established the high-performance core for FITS tensor and table I/O:
- Zero-Copy Tensor I/O: Memory-mapped reads with SIMD-vectorized byte swapping for 1D–4D FITS image extensions.
- Columnar Table Engine: Binary and ASCII table reads with SQL predicate pushdown (
where=) and fast column projection. - PyTorch ML Data Loaders: Native
Datasetclasses (FitsImageDataset,FitsCutoutDataset,FitsCubeDataset,FitsTableDataset) and multi-workermake_loader. - Command-Line Suite: Unix-style CLI tools (
info,header,cutout,convert,verify). - Feature Parity: Comprehensive format support verified against standard FITS test suites.
1.1 — Streaming, Correctness & Remote Hardening (beta)¶
On the same PyTorch ABI lane; the changelog carries the full list:
- Checksum-stamped writes:
write(..., checksum=True)plusverify_checksums. - GIL-free hot reads: DataLoader workers no longer serialize behind one Python thread during disk or network access.
- Auto-adaptive RGB compositing:
transforms.rgb(*bands)andconvert --recipe auto. - Memory-bounded streaming filters:
scan(..., where=...)evaluates predicates per batch, so peak RAM tracksbatch_size. - Multiprocess-safe remote downloads: OS-file-lock dedupe, resumable partials, explicit completeness warnings.
- Silent-corruption fixes: BIT column writes, multi-chunk buffered reads, unsigned-column
where=pushdown.
1.2 — The torch boundary (in progress)¶
Metadata has never needed PyTorch, but every metadata call paid about a second
for it, because the single native extension links TORCH_LIBRARIES and imports
torch in its module body. The rule this work establishes — PyTorch loads
only for a call whose documented return type is a torch.Tensor or that takes
device= — is enforced by tests/test_torch_boundary.py and measured by
benchmarks/bench_import_boundary.py.
Landed so far:
import torchfits.hdu883 ms → 2.6 ms andimport torchfits.io882 ms → 34.6 ms, both with torch absent;Header/Cardwork in a torch-free process.- An interpreter-exit defect fixed: the cache hook imported the extension unconditionally, so a process that never loaded it printed a traceback (or a warning) on every exit.
Remaining, in order:
- Torch-free core library. Extract the inspection half of
FITSFile, theSharedReadMetacaches and aparallel_forof our own into a singlelibtorchfits_coreshared library; bind it astorchfits._corewith a build-id guard against_C. Target:read_header1136 ms → ~5 ms, and the metadata CLI commands (info,header,probe,verify,copy,setkey,table) off torch entirely. - Buffer transport for tables. Return owned byte arenas instead of
torch::Tensorbuffers, build Arrow from them zero-copy, and maketable.read_torchonetorch.from_blobdestination. Target:table.read1555 ms → ~140 ms with no torch installed, and one copy fewer on the tensor path. - Image payloads in the core, which also makes
compress/decompresstorch-free — they re-encode bytes and never do pixel math. - Mechanical proof and packaging. A test asserting the core's dynamic dependencies contain no libtorch, docs, and a decision on publishing the core as its own distribution (it helps the metadata and table audience, not the pixel-math commands, which will always need torch).
Current focus¶
- Single-pass arena decode for buffered table reads. Removes the one
remaining significant benchmark deficit vs
fitsio(narrow-table full reads withmmap=False, ~6–17%): decode straight into caller-visible, strided tensors instead of staging whole rows in scratch chunks. An API-visible change targeted at the next minor. - Selective-projection fast path in the same reader, so filtered scans stop paying for whole-row pread when only a few columns are needed.
- Table semantics polish: complex-column dtypes in
schema(), consistent error types across the mutation API. - Native binding signatures:
read_fits_table_filtereddeclares a default oncolumn_namesahead of a requiredfilters, which is legal in C++ but has no Python signature (and the default is dead — nothing can omit it).nb::sig()on the bindings would make the generated stub exact instead of hand-corrected, and would also givehelp()/inspect.signaturereal names where stubgen currently emitsarg0/arg1. - Object-store recipes: row-band caching and range-fetch patterns for S3-style archives on top of the hardened HTTP downloader.
- CLI wave 3: thin
fitsverifyhelper and fpack-style tile controls (no CFITSIO HTTPS drivers — torchfits keeps its own HTTP stack).
Tooling decisions¶
Re-evaluated 2026-09; recorded so they are not silently revisited.
ty(Astral) is deferred to 1.0. Latest is 0.0.80 — pre-1.0 beta, and not a drop-in for mypy (different defaults, different diagnostics). mypy--strictis the gate. Trigger to revisit: the measured cost it would remove — mypy runs cold in ~20 s over 95 source files, which is already tolerable in CI, so the case is ergonomic rather than blocking. Any trial should start as a non-blocking CI job alongside mypy, not a replacement.- nanobind split mode is not applicable. It collapses the wheel matrix to
one wheel per platform by targeting the Python 3.10 stable ABI, but the real
constraint here is
libtorch_python, which is CPython-version-specific, so torchfits would still ship one wheel per Python version. Adopted nanobind 3 for the API/perf improvements only. - Dependency floors are aspirational, not tested. Floors name the oldest
release with wheels for the minimum supported Python (3.10); nothing
installs them. Follow-up that would make them real: a lowest-direct
resolution job (
pip install --resolution lowest-direct, or a pixi minimum env) so the metadata cannot drift from reality again.
2.0.0 — Native C++ / GPU-Direct Architecture (Future)¶
The 2.0 major release aims to drop external legacy C library dependencies in favor of a modern, native C++/CUDA I/O engine:
- Direct Storage-to-GPU Transport (GPUDirect Storage): Direct DMA transfers from NVMe/object storage directly into NVIDIA GPU device memory (cuFile/GDS) without intermediate host-memory bouncing.
- Native Astronomical Tile Codecs: Pure C++20 and CUDA implementations of Rice, H-Compress, and Gzip decompression.
- Asynchronous Batch Execution: Fully non-blocking multi-file decoders scheduled via CUDA streams and CPU worker pools.
- Stable Python API: Maintaining full backwards compatibility with the 1.x
read_tensor,table.read, andtorchfits.dataAPIs.
Scheduled 2.0 removals¶
Deprecated 1.x shims that will be dropped in 2.0:
configure_cpp_cache(deprecated in favor of the cache-manager configuration path).handle_cache_capacity(per-path handle caching was removed; the argument is accepted and ignored).- Legacy
write()dict-vs-tensor argument conventions are documented but will be tightened to explicit keyword-only forms.
Permanent Scope & Design Boundaries¶
To ensure focus, long-term maintainability, and peak performance, torchfits maintains strict scope boundaries:
- No Celestial Coordinate Systems or WCS Math: Coordinate transformations belong in
astropy.wcs.torchfitsoutputs raw pixel tensors with standard header metadata for Astropy consumption. - No Physical Units Engine: Quantity conversions belong in
astropy.units. - No High-Level Astronomy Modeling: Source extraction, PSF fitting, and continuum fitting belong in domain analysis packages (e.g. Photutils, SEP).
- Format Integrity: Strict compliance with the official IAU FITS standard.