API Reference¶
Complete Python API reference for torchfits — covering tensor image I/O, tabular catalog dataframes, multi-extension files, headers, PyTorch dataset loaders, and astronomical stretch transforms.
For step-by-step code tutorials, see Python Workflows and Astronomical Examples.
Which reader?¶
| Goal | Call | Returns | mmap default |
|---|---|---|---|
| Image / cube / spectrum | read_tensor(path, hdu=0) |
torch.Tensor |
True |
| Catalog as Arrow | table.read(path, hdu=1, where=…) |
pyarrow.Table |
True |
| Columns as tensors | table.read_torch(path, hdu=1) |
scalar columns as tensors; VLA columns use list/tuple values | "auto" |
| Polars | table.read_polars(path, hdu=1) |
FITSPolarsFrame wrapper |
(via table.read) |
table.read_arrow is an alias for table.read.
Caches. Clear in-process I/O metadata with clear_file_cache() /
get_cache_performance(). Wipe everything (in-process + disk roots) with
clear_all_caches(). Disk roots and dataset sizing live under
torchfits.cache. See Core I/O → Cache Utilities.
Quick Paths¶
Images and Files¶
| Goal | Entry point | Reference |
|---|---|---|
| Read N-D array as tensor | read_tensor(path, hdu=0, device="cpu", mmap=True) |
Core I/O |
| Read image or table (auto-detect) | read(path, hdu=None, return_header=False) |
Core I/O |
| Rectangular cutout | read_subset(path, hdu, x1, y1, x2, y2) |
Core I/O |
| Repeated cutouts from one file | open_subset_reader(path, hdu=0) |
Core I/O |
| Repeated table column reads | open_table_reader(path, hdu=1) |
Core I/O |
| Multiple HDUs at once | read_hdus(path, hdus=[0, 1, 2]) |
Core I/O |
| Write a tensor | write_tensor(path, tensor, ..., quantize=None) |
Core I/O |
| Read header only | read_header(path, hdu=0) |
Core I/O |
| Table row count (skinny) | read_nrows(path, hdu=1) |
Core I/O |
| Selected header keys (skinny) | read_keys(path, keys, hdu=0) |
Core I/O |
| Image BITPIX+shape (skinny) | read_shape(path, hdu=0) |
Core I/O |
| HDU type / count (skinny) | read_hdu_type / read_num_hdus |
Core I/O |
| EXTNAME of an HDU (skinny) | read_extname(path, hdu=1) |
Core I/O |
| Table colnames / info (skinny) | read_colnames / read_table_info |
Core I/O |
| Multi-HDU context manager | open(path, mode="r") |
Core I/O |
| Batch-read many files | read_batch(file_paths, hdu=0) |
Core I/O |
| Insert an HDU | insert_hdu(path, data, index=1) |
Core I/O |
| Replace an HDU | replace_hdu(path, hdu, data) |
Core I/O |
| Delete an HDU | delete_hdu(path, hdu) |
Core I/O |
| Write FITS checksums | write_checksums(path, hdu=0) |
Core I/O |
| Verify FITS checksums | verify_checksums(path, hdu=0) |
Core I/O |
Tables as dataframes¶
| Goal | Entry point | Reference |
|---|---|---|
| Read dataframe (Arrow) | table.read(path, hdu=1, columns=None, where=None) |
Tables |
| Stream dataframe batches | table.scan(path, hdu=1, batch_size=65536) |
Tables |
| Dataframe columns as tensors | table.read_torch(path, hdu=1, columns=None) |
Tables |
| Stream tensor-column chunks | table.scan_torch(path, hdu=1, batch_size=65536) |
Tables |
| Astropy Table | table.read_astropy(path, hdu=1) |
Tables |
| Native Polars dataframe | table.read_polars(path, hdu=1) |
Tables |
| Streaming Polars batches | table.scan_polars(path, hdu=1) |
Tables |
| DuckDB SQL | table.duckdb_query(path, sql, hdu=1) |
Tables |
Destination-qualified spelling of table.read (same object): table.read_arrow.
| Goal | Entry point | Reference |
|---|---|---|
| Interop: Astropy | to_astropy(table_dict, decode_bytes=False) |
Tables |
| Interop: Polars | to_polars(table_dict, decode_bytes=False) |
Tables |
| Interop: Arrow | to_arrow(table_dict, decode_bytes=False) |
Tables |
| Interop: pandas | to_pandas(table_dict, decode_bytes=False) |
Tables |
Cache¶
| Goal | Entry point | Reference |
|---|---|---|
| Reset in-process + disk caches | clear_all_caches() |
Core I/O |
| Evict one file's cached state | clear_file_cache(data=True, ..., cpp=True) |
Core I/O |
| Cache hit / eviction counters | get_cache_performance() |
Core I/O |
Datasets and loaders¶
Start with Quick start for a first read, then use raw I/O until you need shuffling, workers, or epochs — see Data module and Transforms. Runnable scripts are indexed in Examples.
| Goal | Entry point | Reference |
|---|---|---|
| General N-D image, map-style | FitsTensorDataset(paths, hdu=0, label_key=None) |
Data |
| General N-D image, iterable (multi-worker) | FitsTensorIterableDataset(paths, shuffle=False) |
Data |
| 2D image peer | FitsImageDataset(paths, hdu=0) |
Data |
| 2D image iterable peer | FitsImageIterableDataset(paths, shuffle=False) |
Data |
| 3D+ cube peer | FitsCubeDataset(paths, hdu=0, slice_index=None) |
Data |
| 3D+ cube iterable peer | FitsCubeIterableDataset(paths, hdu=0, slice_index=None) |
Data |
| 1D / multi-arm spectrum (map-style) | FitsSpectrumDataset(paths, hdu=0, layout="dict") |
Data |
| 1D / multi-arm spectrum (iterable) | FitsSpectrumIterableDataset(paths, hdu=0, layout="dict") |
Data |
| Table map-style (fits in RAM) | FitsTableDataset(path, hdu=1) |
Data |
| Table streaming (large) | FitsTableIterableDataset(path, hdu=1, batch_size=65536) |
Data |
| Cutout patches | FitsCutoutDataset(cutouts) |
Data |
| Survey mosaic staged cutouts | FitsStagedCutoutIterableDataset(paths, cutouts_per_file=500) |
Data |
| DataLoader + cache defaults | make_loader(dataset, batch_size=32) |
Data |
| Image stretches, normalizers, clip | from torchfits.transforms import … |
Transforms |
| Shell inspect / convert | torchfits info|header|verify|… |
CLI |
Reference Pages¶
| Page | What it covers |
|---|---|
| CLI | torchfits command-line tools, exit codes, MEF defaults |
| Core I/O | read, read_tensor, read_subset, read_hdus, write_tensor, write, open, headers, HDU mutation, checksums, batch reads, cache |
| Tables | FITS tables as dataframes: table.read / read_torch / read_polars, mutations, interop |
| Data | FitsTensorDataset, FitsTensorIterableDataset, FitsCubeDataset, FitsSpectrumDataset, table/cutout datasets, make_loader, remote prefetch |
| Transforms | Transform classes (callable protocol, not nn.Module) with verified math, parameters, invertibility, and when-to-use guidance |
| Architecture | C++/Python layering, I/O paths, caching, threading, CFITSIO mapping, environment variables |
Package Namespaces¶
| Namespace | Purpose |
|---|---|
torchfits (root) |
I/O functions and HDU classes |
torchfits.hdu |
HDU/header types (Header, Card, HDUList, …) |
torchfits.table |
FITS table / dataframe I/O, mutation, interop |
torchfits.data |
Dataset classes and loader factory |
torchfits.transforms |
Transform classes |
torchfits.cache |
Cache configuration and management |
torchfits.where |
Predicate parser and evaluator |
torchfits.cpp |
Low-level native compatibility surface |
torchfits.cpp is the low-level native compatibility surface used by
performance-sensitive downstream packages. Its __all__ is the
function-level compatibility contract; new compiled-extension symbols are
private until promoted there.
Public API Index¶
The root torchfits.__all__ (public surface; see __all__):
- Image / table I/O:
read,write,open,read_tensor,read_subset,read_hdus,read_batch,read_batch_info,open_subset_reader,open_table_reader,write_tensor - Skinny metadata:
read_header,read_colnames,read_extname,read_hdu_type,read_keys,read_nrows,read_num_hdus,read_shape,read_table_info - Checksums / HDU mutation:
verify_checksums,write_checksums,insert_hdu,replace_hdu,delete_hdu - Cache:
get_cache_performance,clear_all_caches,clear_file_cache - Interop:
to_astropy,to_pandas,to_arrow,to_polars - HDU types:
Header,Card,HDUList,TensorHDU,TableHDU,TableHDURef - Namespaces:
table,cache,transforms,data,where,hdu
Each entry has a quick-path row and a reference page above.
HDU Types¶
| Class | Description |
|---|---|
TensorHDU |
Image HDU with lazy .data (returns DataView) and .header |
TableHDU |
In-memory table HDU with tensor columns |
TableHDURef |
Lazy file-backed table handle |
Header |
Dict-like FITS header preserving card order and semantics |
Card |
Single FITS header card: Card(key, value, comment) |
Limitations¶
- VLA columns use buffered I/O; mmap reads and in-place updates are not supported.
- Scaled table columns do not support mmap updates; use the buffered path.
- Non-CPU image tensors are copied to host before FITS writes; table-column handling has a separate backend path.
- Compressed writes accept tensor-convertible IMAGE-HDU payloads; table mappings are also supported, but dict image descriptors must contain tensor-convertible data.