Skip to content
torchfits

High-performance FITS I/O

Read and write FITS as tensors and tables — cutouts, filters, and a shell CLI.

import torchfits
tensor = torchfits.read_tensor("science.fits", device="cuda")
df     = torchfits.table.read("catalog.fits", where="MAG < 20")

Home

torchfits delivers high-performance FITS I/O for the modern Python data science and machine learning ecosystem. Powered by a native C++ engine with vendored CFITSIO, it decodes astronomical images directly to PyTorch GPU/CPU tensors, queries binary catalogs as Arrow tables with C++ predicate pushdown, reads survey cutouts without materializing full frames, and provides a full shell CLI.

Prebuilt binary wheels are available for Linux (x86_64, aarch64) and macOS (Apple Silicon arm64) supporting Python 3.10–3.14 (no C++ compiler or system libraries needed).


Core Capabilities

  • C++ Engine & Vendored CFITSIO


    Decodes FITS structures at near-native C speed. Ships with CFITSIO vendored in prebuilt wheels — install and run with zero compiler setup.

    Installation Guide

  • GPU Tensor Placement


    Decode on the host, then place the result on CUDA or Apple Silicon unified memory in one call (device="cuda" / "mps").

    Core I/O Workflows

  • Arrow-Native Catalogs & SQL


    Filter multi-million row catalogs with C++ pushdown predicates (where="MAG < 20"). Zero-copy where dtypes allow; integration with Polars, Pandas, and DuckDB.

    Table Workflows

  • Survey-Scale Cutouts & ML


    Extract sub-regions and postage stamps from giant mosaics without RAM exhaustion. Native PyTorch Dataset and multi-worker make_loader.

    Machine Learning with FITS


At a Glance

import torchfits

# Read image directly onto GPU or CPU
tensor = torchfits.read_tensor("science.fits", hdu="SCI", device="cuda")
print(f"Loaded tensor on {tensor.device} with shape {tensor.shape}")

# Read pixel data alongside header cards
data, header = torchfits.read("science.fits", return_header=True)
print("Object:", header.get("OBJECT"))
import torchfits

# Fast C++ predicate pushdown into PyArrow Table
table = torchfits.table.read(
    "catalog.fits",
    hdu=1,
    columns=["RA", "DEC", "MAG_G"],
    where="MAG_G < 20.0 AND CLASS_STAR > 0.8",
)

# Convert to Polars or Pandas without data copies
import polars as pl

df = pl.from_arrow(table)
import torchfits

# Single bounding box cutout [x1, y1, x2, y2)
stamp = torchfits.read_subset("mosaic.fits", hdu=0, x1=100, y1=100, x2=228, y2=228)

# High-throughput batch cutouts using reusable file handle
with torchfits.open_subset_reader("survey_mosaic.fits", hdu=0) as reader:
    stamp_a = reader.read_subset(100, 100, 228, 228)
    stamp_b = reader.read_subset(500, 500, 628, 628)
# Inspect extensions, shapes, and data types
torchfits info science.fits

# Dump header keywords or build cross-file summary catalogs
torchfits header *.fits --keyword-table -k OBJECT -k FILTER

# Filter binary tables and export directly to Apache Parquet
torchfits convert catalog.fits bright.parquet -e 1 -w "MAG_G < 18.0"

Why torchfits?

Capability Astropy (io.fits) / fitsio torchfits
GPU Tensor Placement Manual host read \(\rightarrow\) .to(device) copy Native device="cuda" / "mps" decode
Catalog Query & Slicing Load full table to Python \(\rightarrow\) boolean mask C++ in-engine where= pushdown to PyArrow
Mosaic Cutouts Section indexing or full-frame slicing Zero-overhead read_subset & open_subset_reader
Machine Learning Pipelines Custom boilerplate wrapper classes Native FitsImageDataset + multi-worker make_loader
Command-Line Suite Disparate tools (fitsinfo, imstat, fpack) Unified, multi-core torchfits CLI suite
Packaging & Installation May require local C compilation Prebuilt wheels with vendored CFITSIO

Explore detailed benchmarks and published results in Benchmarks, or check feature coverage in the Compatibility & Parity Matrix.


Documentation Pathways


Docs channels: stable (latest release) · edge (main tip).