Skip to content

API Reference

Complete Python API reference for torchfits — covering tensor image I/O, tabular catalog dataframes, multi-extension files, headers, PyTorch dataset loaders, and astronomical stretch transforms.

For step-by-step code tutorials, see Python Workflows and Astronomical Examples.


Which reader?

Goal Call Returns mmap default
Image / cube / spectrum read_tensor(path, hdu=0) torch.Tensor True
Catalog as Arrow table.read(path, hdu=1, where=…) pyarrow.Table True
Columns as tensors table.read_torch(path, hdu=1) scalar columns as tensors; VLA columns use list/tuple values "auto"
Polars table.read_polars(path, hdu=1) FITSPolarsFrame wrapper (via table.read)

table.read_arrow is an alias for table.read.

Caches. Clear in-process I/O metadata with clear_file_cache() / get_cache_performance(). Wipe everything (in-process + disk roots) with clear_all_caches(). Disk roots and dataset sizing live under torchfits.cache. See Core I/O → Cache Utilities.


Quick Paths

Images and Files

Goal Entry point Reference
Read N-D array as tensor read_tensor(path, hdu=0, device="cpu", mmap=True) Core I/O
Read image or table (auto-detect) read(path, hdu=None, return_header=False) Core I/O
Rectangular cutout read_subset(path, hdu, x1, y1, x2, y2) Core I/O
Repeated cutouts from one file open_subset_reader(path, hdu=0) Core I/O
Repeated table column reads open_table_reader(path, hdu=1) Core I/O
Multiple HDUs at once read_hdus(path, hdus=[0, 1, 2]) Core I/O
Write a tensor write_tensor(path, tensor, ..., quantize=None) Core I/O
Read header only read_header(path, hdu=0) Core I/O
Table row count (skinny) read_nrows(path, hdu=1) Core I/O
Selected header keys (skinny) read_keys(path, keys, hdu=0) Core I/O
Image BITPIX+shape (skinny) read_shape(path, hdu=0) Core I/O
HDU type / count (skinny) read_hdu_type / read_num_hdus Core I/O
EXTNAME of an HDU (skinny) read_extname(path, hdu=1) Core I/O
Table colnames / info (skinny) read_colnames / read_table_info Core I/O
Multi-HDU context manager open(path, mode="r") Core I/O
Batch-read many files read_batch(file_paths, hdu=0) Core I/O
Insert an HDU insert_hdu(path, data, index=1) Core I/O
Replace an HDU replace_hdu(path, hdu, data) Core I/O
Delete an HDU delete_hdu(path, hdu) Core I/O
Write FITS checksums write_checksums(path, hdu=0) Core I/O
Verify FITS checksums verify_checksums(path, hdu=0) Core I/O

Tables as dataframes

Goal Entry point Reference
Read dataframe (Arrow) table.read(path, hdu=1, columns=None, where=None) Tables
Stream dataframe batches table.scan(path, hdu=1, batch_size=65536) Tables
Dataframe columns as tensors table.read_torch(path, hdu=1, columns=None) Tables
Stream tensor-column chunks table.scan_torch(path, hdu=1, batch_size=65536) Tables
Astropy Table table.read_astropy(path, hdu=1) Tables
Native Polars dataframe table.read_polars(path, hdu=1) Tables
Streaming Polars batches table.scan_polars(path, hdu=1) Tables
DuckDB SQL table.duckdb_query(path, sql, hdu=1) Tables

Destination-qualified spelling of table.read (same object): table.read_arrow.

Goal Entry point Reference
Interop: Astropy to_astropy(table_dict, decode_bytes=False) Tables
Interop: Polars to_polars(table_dict, decode_bytes=False) Tables
Interop: Arrow to_arrow(table_dict, decode_bytes=False) Tables
Interop: pandas to_pandas(table_dict, decode_bytes=False) Tables

Cache

Goal Entry point Reference
Reset in-process + disk caches clear_all_caches() Core I/O
Evict one file's cached state clear_file_cache(data=True, ..., cpp=True) Core I/O
Cache hit / eviction counters get_cache_performance() Core I/O

Datasets and loaders

Start with Quick start for a first read, then use raw I/O until you need shuffling, workers, or epochs — see Data module and Transforms. Runnable scripts are indexed in Examples.

Goal Entry point Reference
General N-D image, map-style FitsTensorDataset(paths, hdu=0, label_key=None) Data
General N-D image, iterable (multi-worker) FitsTensorIterableDataset(paths, shuffle=False) Data
2D image peer FitsImageDataset(paths, hdu=0) Data
2D image iterable peer FitsImageIterableDataset(paths, shuffle=False) Data
3D+ cube peer FitsCubeDataset(paths, hdu=0, slice_index=None) Data
3D+ cube iterable peer FitsCubeIterableDataset(paths, hdu=0, slice_index=None) Data
1D / multi-arm spectrum (map-style) FitsSpectrumDataset(paths, hdu=0, layout="dict") Data
1D / multi-arm spectrum (iterable) FitsSpectrumIterableDataset(paths, hdu=0, layout="dict") Data
Table map-style (fits in RAM) FitsTableDataset(path, hdu=1) Data
Table streaming (large) FitsTableIterableDataset(path, hdu=1, batch_size=65536) Data
Cutout patches FitsCutoutDataset(cutouts) Data
Survey mosaic staged cutouts FitsStagedCutoutIterableDataset(paths, cutouts_per_file=500) Data
DataLoader + cache defaults make_loader(dataset, batch_size=32) Data
Image stretches, normalizers, clip from torchfits.transforms import … Transforms
Shell inspect / convert torchfits info|header|verify|… CLI

Reference Pages

Page What it covers
CLI torchfits command-line tools, exit codes, MEF defaults
Core I/O read, read_tensor, read_subset, read_hdus, write_tensor, write, open, headers, HDU mutation, checksums, batch reads, cache
Tables FITS tables as dataframes: table.read / read_torch / read_polars, mutations, interop
Data FitsTensorDataset, FitsTensorIterableDataset, FitsCubeDataset, FitsSpectrumDataset, table/cutout datasets, make_loader, remote prefetch
Transforms Transform classes (callable protocol, not nn.Module) with verified math, parameters, invertibility, and when-to-use guidance
Architecture C++/Python layering, I/O paths, caching, threading, CFITSIO mapping, environment variables

Package Namespaces

Namespace Purpose
torchfits (root) I/O functions and HDU classes
torchfits.hdu HDU/header types (Header, Card, HDUList, …)
torchfits.table FITS table / dataframe I/O, mutation, interop
torchfits.data Dataset classes and loader factory
torchfits.transforms Transform classes
torchfits.cache Cache configuration and management
torchfits.where Predicate parser and evaluator
torchfits.cpp Low-level native compatibility surface

torchfits.cpp is the low-level native compatibility surface used by performance-sensitive downstream packages. Its __all__ is the function-level compatibility contract; new compiled-extension symbols are private until promoted there.


Public API Index

The root torchfits.__all__ (public surface; see __all__):

  • Image / table I/O: read, write, open, read_tensor, read_subset, read_hdus, read_batch, read_batch_info, open_subset_reader, open_table_reader, write_tensor
  • Skinny metadata: read_header, read_colnames, read_extname, read_hdu_type, read_keys, read_nrows, read_num_hdus, read_shape, read_table_info
  • Checksums / HDU mutation: verify_checksums, write_checksums, insert_hdu, replace_hdu, delete_hdu
  • Cache: get_cache_performance, clear_all_caches, clear_file_cache
  • Interop: to_astropy, to_pandas, to_arrow, to_polars
  • HDU types: Header, Card, HDUList, TensorHDU, TableHDU, TableHDURef
  • Namespaces: table, cache, transforms, data, where, hdu

Each entry has a quick-path row and a reference page above.


HDU Types

Class Description
TensorHDU Image HDU with lazy .data (returns DataView) and .header
TableHDU In-memory table HDU with tensor columns
TableHDURef Lazy file-backed table handle
Header Dict-like FITS header preserving card order and semantics
Card Single FITS header card: Card(key, value, comment)

Limitations

  • VLA columns use buffered I/O; mmap reads and in-place updates are not supported.
  • Scaled table columns do not support mmap updates; use the buffered path.
  • Non-CPU image tensors are copied to host before FITS writes; table-column handling has a separate backend path.
  • Compressed writes accept tensor-convertible IMAGE-HDU payloads; table mappings are also supported, but dict image descriptors must contain tensor-convertible data.