Benchmarks¶
Headline (v1.1.0, CANFAR Linux, extended
labtiming profile, mmap on+off matrix; git-mirroredexhaustive_cpu_20260807_013736/exhaustive_cuda_20260807_013736): torchfits wins 100% of significant image comparisons on both CPU and CUDA hosts — compressed and uncompressed, every dtype, every transport — and 98.9–99.5% of smart-family table comparisons. The remaining significant case family where a peer is ahead is narrow-table full reads withmmap=False(fitsio, by 21–36% on CPU and 8% on CUDA): fitsio reads one column at a time while our buffered path stages whole rows; the structural fix (single-pass decode into caller-visible memory) lands in 1.2. Image HCOMPRESS lags vs fitsio are sub-1.03× noise on the shared-CFITSIO path. Full per-cell data below; CSVs under Published CSVs.
torchfits benchmarks cover FITS tensor I/O (IMAGE HDUs, typically 1D–4D)
and FITS table I/O vs Astropy and fitsio across CPU and GPU hardware.
A note on fairness: Headline ratios below are medians from reproducible benchmark runs across our test suites (the lab timing profile uses more warmup and repetitions than the quick user profile) — not guarantees on your specific filesystem, file mix, or PyTorch version. Check Performance comparisons & limitations for a transparent breakdown of cases where peer libraries are competitive or faster.
How to read this page¶
| If you want… | Jump to |
|---|---|
| Headline wins | Performance highlights |
| Cases where torchfits is not #1 (CPU and GPU) | Performance comparisons & limitations |
| GPU transport rows | I/O transport and backend |
| Python × PyTorch version variance | Version matrix variance |
| Reproduce numbers | Reproducing |
| Every measured configuration | Exhaustive benchmark results |
| Raw CSV | Published CSVs |
Published GPU/CPU numbers come from the multi-host benchmark runs
(exhaustive_mps_*, exhaustive_cpu_*, exhaustive_cuda_*). Manual
workflow_dispatch on .github/workflows/bench-report.yml is CPU-only and
does not refresh GPU cells.
Comparison targets¶
| Domain | torchfits API | Compared against |
|---|---|---|
| Tensor (IMAGE HDU) | read / read_tensor / write |
astropy.io.fits, fitsio |
| Table (dataframe) | torchfits.table |
astropy.io.fits, fitsio |
Methodology¶
Each case measures median wall-clock time over multiple repetitions, plus peak process RSS (and peak CUDA alloc when on CUDA). Performance ranking is time-based; RSS is reported alongside times.
Cases are grouped into two families:
- default — high-level API (
torchfits.read/table.read, etc.). - specialized —
torchfits_specializedmethods (open-once handle /open_subset_readerpaths). Empty specialized cells mean that path was not measured for the case.
Fairness controls:
- Rows with mismatched mmap behavior are marked
SKIPPEDand excluded from rankings. - Why fitsio has no mmap rows: fitsio does not expose a comparable mmap
toggle. Under
mmap_target=on/strict_mmap_fairness, fitsio rows are non-comparable and show as skipped in transport tables (seescripts/render_bench_iopath_table.py). - Warm-cache and cold-cache profiles are kept separate.
Disk to GPU¶
True disk→GPU (GPUDirect Storage / cuFile, or a CFITSIO path that never
touches host RAM) is not implemented. Every Python FITS stack here
decodes on the host, then copies with .to(device). Exploring a direct path
is a 2.0 item (see Roadmap) — not a 1.x claim.
Tables on GPU transports¶
Table GPU transport rows compare table.read_torch(..., device=cpu) against
device=cuda / device=mps on a medium mixed catalog case
(mixed_100000). Decode still happens on the host; the GPU column measures
host decode plus H2D copy into tensor columns.
CUDA Host-to-Device Transfer & Small Payloads¶
For small tensor payloads (e.g. 1D arrays and small \(64 \times 64\) sub-regions), fixed kernel launch latency and Host-to-Device (H2D) memory transfer dominate over raw decode throughput.
In torchfits, the C++ engine optimizes memory transfers by coordinating host buffers and asynchronous CUDA streams:
- For larger images, direct memory transfers match peak PCIe bus bandwidth.
- For small payloads, latency remains competitive with in-memory transfers, operating at parity with baseline libraries on NVIDIA CUDA and Apple Silicon MPS.
Vectorized SIMD Integer Decoding (Unsigned Integers & BZERO)¶
Standard astronomical FITS stores unsigned 16-bit and 32-bit integers using signed formats paired with standard BZERO offsets (\(y = \text{raw} + 32768\)).
torchfits fuses big-endian byte-swapping and BZERO offset calculations directly into vectorized SIMD loops within the C++ engine:
- Eliminates secondary scalar normalization passes over memory.
- Delivers up to \(3\times\) speedups on large uint16 and uint32 image arrays compared to two-stage Python conversions.
Python & PyTorch Matrix Variance¶
The headline benchmarks on this page are reported using the Pareto-optimal (champion) environment combination measured across our CANFAR exhaustive matrix: PyTorch 2.12 + Python 3.11.
Below is the measured performance variance across the full matrix grid (Python 3.10–3.14 × PyTorch 2.10–2.13 × CPU / CUDA), tracking the average latency delta and overhead relative to the champion configuration.
Summary: Average Performance Overhead vs Champion¶
- CPU Host Workloads:
- Optimal Baseline: PyTorch 2.12 + Python 3.11 (Geometric-mean latency: 0.106 ms)
- Average Python Variance: Across all Python versions (3.10–3.14), average latency penalty is +9.5% (+7.6% on 3.10, +11.2% on 3.11, +10.5% on 3.12, +8.4% on 3.13, +12.8% on 3.14).
-
Average PyTorch Variance: Across PyTorch minor versions (2.10–2.13), average latency penalty is +9.8% (+10.7% on 2.10, +12.3% on 2.11, +7.4% on 2.12, +9.9% on 2.13).
-
CUDA Workloads (NVIDIA GPU):
- Optimal Baseline: PyTorch 2.12 + Python 3.11 (Geometric-mean latency: 0.187 ms)
- Average Python Variance: Across all Python versions (3.10–3.14), average latency penalty is +5.8% (+8.6% on 3.10, +4.0% on 3.11, +5.0% on 3.12, +4.9% on 3.13, +6.3% on 3.14).
- Average PyTorch Variance: Across PyTorch minor versions (2.10–2.13), average latency penalty is +5.7% (+3.7% on 2.10, +4.3% on 2.11, +3.9% on 2.12, +11.0% on 2.13).
Full Matrix Benchmark Comparison¶
CPU Transport Matrix¶
| PyTorch | Python | Device | Geom Mean (ms) | Delta vs Best | Relative Perf |
|---|---|---|---|---|---|
| 2.12 | 3.11 | CPU | 0.106 | Baseline (Best) | 1.00× |
| 2.10 | 3.10 | CPU | 0.108 | +2.0% slower | 1.02× |
| 2.13 | 3.12 | CPU | 0.111 | +5.0% slower | 1.05× |
| 2.11 | 3.10 | CPU | 0.113 | +6.4% slower | 1.06× |
| 2.11 | 3.13 | CPU | 0.113 | +6.4% slower | 1.06× |
| 2.12 | 3.13 | CPU | 0.113 | +6.6% slower | 1.07× |
| 2.12 | 3.10 | CPU | 0.114 | +7.3% slower | 1.07× |
| 2.13 | 3.13 | CPU | 0.114 | +7.3% slower | 1.07× |
| 2.12 | 3.14 | CPU | 0.114 | +7.4% slower | 1.07× |
| 2.13 | 3.14 | CPU | 0.115 | +8.6% slower | 1.09× |
| 2.10 | 3.12 | CPU | 0.116 | +8.9% slower | 1.09× |
| 2.10 | 3.14 | CPU | 0.117 | +10.2% slower | 1.10× |
| 2.11 | 3.11 | CPU | 0.119 | +11.7% slower | 1.12× |
| 2.11 | 3.12 | CPU | 0.119 | +11.9% slower | 1.12× |
| 2.10 | 3.13 | CPU | 0.120 | +13.2% slower | 1.13× |
| 2.13 | 3.11 | CPU | 0.121 | +13.9% slower | 1.14× |
| 2.13 | 3.10 | CPU | 0.122 | +14.9% slower | 1.15× |
| 2.12 | 3.12 | CPU | 0.123 | +16.0% slower | 1.16× |
| 2.10 | 3.11 | CPU | 0.126 | +19.2% slower | 1.19× |
| 2.11 | 3.14 | CPU | 0.132 | +24.9% slower | 1.25× |
CUDA Transport Matrix¶
| PyTorch | Python | Device | Geom Mean (ms) | Delta vs Best | Relative Perf |
|---|---|---|---|---|---|
| 2.12 | 3.11 | CUDA | 0.187 | Baseline (Best) | 1.00× |
| 2.10 | 3.14 | CUDA | 0.189 | +1.1% slower | 1.01× |
| 2.10 | 3.13 | CUDA | 0.190 | +1.8% slower | 1.02× |
| 2.10 | 3.12 | CUDA | 0.190 | +1.8% slower | 1.02× |
| 2.12 | 3.13 | CUDA | 0.191 | +1.9% slower | 1.02× |
| 2.11 | 3.12 | CUDA | 0.193 | +3.1% slower | 1.03× |
| 2.11 | 3.14 | CUDA | 0.193 | +3.3% slower | 1.03× |
| 2.12 | 3.12 | CUDA | 0.195 | +4.1% slower | 1.04× |
| 2.11 | 3.11 | CUDA | 0.196 | +5.0% slower | 1.05× |
| 2.11 | 3.13 | CUDA | 0.197 | +5.1% slower | 1.05× |
| 2.10 | 3.11 | CUDA | 0.197 | +5.2% slower | 1.05× |
| 2.11 | 3.10 | CUDA | 0.197 | +5.2% slower | 1.05× |
| 2.12 | 3.10 | CUDA | 0.197 | +5.4% slower | 1.05× |
| 2.13 | 3.11 | CUDA | 0.198 | +5.8% slower | 1.06× |
| 2.12 | 3.14 | CUDA | 0.202 | +8.1% slower | 1.08× |
| 2.10 | 3.10 | CUDA | 0.203 | +8.8% slower | 1.09× |
| 2.13 | 3.12 | CUDA | 0.207 | +10.8% slower | 1.11× |
| 2.13 | 3.13 | CUDA | 0.207 | +10.9% slower | 1.11× |
| 2.13 | 3.14 | CUDA | 0.211 | +12.6% slower | 1.13× |
| 2.13 | 3.10 | CUDA | 0.215 | +15.1% slower | 1.15× |
Published Benchmark Data¶
Exhaustive benchmark datasets and analysis CSVs (results.csv, torchfits_deficits.csv) are published with each release and mirrored under docs/assets/bench/<run-id>/:
exhaustive_cpu_20260807_013736/results.csvexhaustive_cuda_20260807_013736/results.csvexhaustive_mps_20260719_143706/results.csvexhaustive_cpu_20260719_144337/results.csvexhaustive_cuda_20260719_144457/results.csv
Modular Suites & Release Exhaustives¶
Named suites live in benchmarks/suites.py and resolve to bench_all.py flags
(--scope / --filter / --operation / GPU / mmap / profile):
pixi run bench-suite hcompress
pixi run bench-suite compressed_rice -- --no-mmap
pixi run bench-suite fitstable_predicate
pixi run bench-deficit-focus # registry-driven deficit clusters
Release composition is the release suite (full fits + fitstable, mmap matrix,
GPU when present). Host recipes:
| Task | Host | Run ID prefix |
|---|---|---|
pixi run bench-exhaustive-local |
Mac CPU + MPS | exhaustive_mps_* |
pixi run bench-exhaustive-canfar-cpu |
CANFAR multicore CPU | exhaustive_cpu_* |
pixi run bench-exhaustive-canfar-cuda |
CANFAR CUDA | exhaustive_cuda_* |
pixi run bench-release-scorecard -- <run_dir>... |
meta | patches multi-host docs |
pixi run bench-cfitsio-direct |
local C | full-suite pure vendored CFITSIO (--profile full) |
pixi run bench-megacam |
local | CFHT MegaCam MEF cutouts (requires fetched sample data) |
pixi run bench-ml |
local | PyTorch DataLoader throughput vs fitsio |
CFHT MegaCam Cutout Suite¶
Public CFHT MegaCam MEF samples (CADC Direct Data Service) exercise Rice
.fz repeated cutouts with peer ranking:
| Method | Role |
|---|---|
torchfits_cached / fitsio_cached |
Open once + N× subset (comparable family) |
torchfits_materialize |
Decompress plane once, then host slices (isolates Rice vs cutout API) |
torchfits_naive |
Re-open per cutout (pathological baseline; not ranked) |
Uses ZNAXIS* for tile-compressed sizes; throughput is cutout payload MB/s.
bash scripts/fetch_cfht_megacam_sample.sh # once; idempotent
pixi run bench-megacam
Outputs land in benchmarks_results/<run-id>/megacam_results.csv.
On multi-extension CFHT MegaCam exposures (40 cutouts \(\times 256 \times 256\) per HDU), torchfits_cached outperforms fitsio_cached by 7.5%–15.2% across sampled HDUs due to optimized tile decompression handles.
For uncompressed survey mosaics (e.g. CFHTLS MegaPipe float32 stacks),
open_subset_reader maps the data segment once and slices cutouts with
endian swap into torch tensors — see
ML with FITS. Rice .fz
MegaCam cutouts remain a separate comparison (tile decompress inside CFITSIO).
Cold-start and the torch boundary¶
Operation-level benchmarks measure a call; they cannot show what a process
costs to start. That is where PyTorch lives: a header peek reads a 2880-byte
block in microseconds, yet every metadata entry point used to pay for an
image-size tensor runtime it never touched, because the single native extension
links TORCH_LIBRARIES and imports torch in its module body.
The rule this project holds to: PyTorch is loaded at exactly one boundary —
the first call whose documented return type is a torch.Tensor, or that takes
device=. Nothing before it may import torch. tests/test_torch_boundary.py
enforces that in fresh interpreters with import torch blocked outright, and
benchmarks/bench_import_boundary.py records what it costs.
Baseline (cold process, spawn to exit, minimum of 3, macOS arm64, 2026-09). "torch" marks entry points that must not load it:
| Entry point | Cold ms | Loads torch | Budget |
|---|---|---|---|
import torchfits |
49 | no | 250 ms |
import torchfits.hdu |
1113 | yes | 250 ms |
import torchfits.io |
1128 | yes | 250 ms |
import torchfits.table |
1132 | yes | 500 ms |
read_header |
1136 | yes | 250 ms |
read_keys |
1131 | yes | 250 ms |
read_colnames |
1118 | yes | 250 ms |
read_num_hdus |
1135 | yes | 250 ms |
read_shape |
1125 | yes | 250 ms |
read_table_info |
1132 | yes | 250 ms |
open + hdul[1].header |
1132 | yes | 250 ms |
table.read (Arrow) |
1555 | yes | 600 ms |
table.schema |
1422 | yes | 600 ms |
read_tensor (tensor destination) |
1134 | yes | — |
table.read_torch (tensor destination) |
1150 | yes | — |
import torch (reference) |
1114 | yes | — |
The ~49 ms floor is interpreter start; the remaining ~1070 ms is PyTorch. Every row above the tensor-destination pair is paying it for a header read. The harness writes its own minimal FITS file with the standard library, so it runs in a torch-free environment either way:
pixi run python benchmarks/bench_import_boundary.py # report
pixi run python benchmarks/bench_import_boundary.py --strict # gate
Correctness checks¶
| Check | Command | Validates |
|---|---|---|
| fitsio parity | pixi run pytest tests/test_fitsio_upstream_smoke.py -q |
Common fitsio image, header, table, compression, and checksum workflows |
| Astropy parity | pixi run pytest tests/test_astropy_upstream_smoke.py -q |
Common Astropy HDU, header, image, compressed-image, table, and scaled-data workflows |
| Package isolation | pixi run pytest tests/test_package_isolation.py tests/test_docs_integrity.py -q |
Clean FITS-only package boundary and docs contract |
Reproducing¶
pixi run bench-fits
pixi run bench-fitstable
pixi run bench-all
pixi run bench-ml
bash scripts/fetch_cfht_megacam_sample.sh && pixi run bench-megacam
# Full transport matrix (mmap on + off, doubles CPU rows; GPU rows for both when CUDA/MPS):
pixi run -e bench-gpu python benchmarks/bench_all.py --profile lab --scope all --mmap-matrix
For focused FITS partitions:
pixi run -e bench-all python benchmarks/bench_all.py --scope fits --filter '^(tiny_)'
pixi run -e bench-all python benchmarks/bench_all.py --scope fits --filter '^(small_)'
pixi run -e bench-all python benchmarks/bench_all.py --scope fits --filter '^(medium_|large_)'
pixi run -e bench-all python benchmarks/bench_all.py --scope fits --filter '^(scaled_|compressed_|mef_)'
Named focused-benchmark recipes (mmap on+off, no unrelated GPU matrix when scoped to tables):
pixi run bench-deficit-focus # hcompress + tiny_int8 + narrow predicates
pixi run bench-deficit-focus hcompress
pixi run bench-deficit-focus tiny_int8
pixi run bench-deficit-focus predicate
Rankings and comparisons group by (domain, case_id, family, mmap_target) so
mmap-on and mmap-off peers are never cross-compared.
Benchmark Scripts¶
| Script | Domain | Description |
|---|---|---|
bench_all.py |
fits / fitstable | FITS benchmark orchestrator |
bench_fits_io.py |
fits | Image I/O across dtypes, sizes, compression, scaling, MEF, and cutouts |
bench_fitstable_io.py |
fitstable | Table I/O across row counts, schemas, projection, row slicing, predicates, and streaming |
bench_all.py / bench-fits |
fits | Published-results path |
bench_table.py |
fitstable | Table API timing |
bench_arrow_tables.py |
fitstable | Arrow-oriented table workflows |
bench_gpu_transports.py |
fits (GPU) | CUDA/MPS image reads, cutouts, repeated cutouts (disk→CPU→GPU / disk→RAM→GPU rows) |
bench_ml_loader.py |
fits (diagnostic) | PyTorch DataLoader throughput (not merged into bench-all CSV) |
bench_gpu_memory.py |
fits (diagnostic) | GPU memory/leak checks (non-gating) |
bench_denoise.py |
ml (scientific) | Noise2Noise CR-cleaning on real CFHT MegaCam frames (dark→blank framing, torchfits loaders vs Astropy; see denoise-pipeline.md) |
bench_import_boundary.py |
cold start | Fresh-process spawn-to-exit cost per entry point, split by whether torch was loaded; --strict gates on the boundary and per-entry-point budgets |
Coverage matrix¶
What the exhaustive bench-all suite measures today, and what is intentionally out of
scope or not yet wired into the published tables.
| Dimension | Covered? | Where | Gap / caveat |
|---|---|---|---|
| Backends (torchfits / astropy / fitsio) | Yes | bench_fits_io.py, bench_fitstable_io.py |
fitsio often excluded from mmap-fairness summaries; uint image comparators may be torchfits-only when astropy requires buffered fallback |
| CPU vs GPU device | Partial | CPU: full matrix; GPU: tensor reads | GPU requires CUDA/MPS (pixi run -e bench-gpu); manual CI bench is CPU-only |
I/O transport disk→RAM→CPU |
Yes | bench-all mmap-on pass |
Median mixes many ops/sizes — coarse aggregate |
I/O transport disk→CPU (non-mmap) |
Yes | bench-all --mmap-matrix mmap-off pass |
Buffered host decode |
I/O transport disk→RAM→GPU |
Partial | bench_gpu_transports.py (mmap on) |
Tensor read_full, cutouts, repeated cutouts; tables until suite lands |
I/O transport disk→CPU→GPU |
Partial | bench_gpu_transports.py (mmap off) |
Same with buffered host decode + H2D |
I/O transport disk→GPU |
No | — | No host-bypass path yet (see Methodology); 2.0 / roadmap |
| BITPIX / dtypes | Partial | int8–int64, float32/64 × 1D/2D/3D | Native uint16/uint32 2D sample datasets; unsigned via BZERO in scaled_* |
| Tensor dimensions / sizes | Yes | tiny → large; 1D–3D (4D where sample datasets exist) | Large 3D cubes may hit size caps |
| Compression (read) | Yes | gzip, rice, hcompress, plio | Write→compress cases are being added to the suite |
| Scaling (BSCALE/BZERO) | Yes | scaled_small/medium/large |
Table-column scaling not isolated |
| Random / repeated access | Yes | cutouts, random_ext_full_reads_200, open_subset_reader |
MEF random ext reads on selected sample datasets |
| Multi-extension (MEF) | Yes | mef_*, multi_mef_10ext, MegaCam suite |
— |
| Table full read / projection / slice | Yes | bench_fitstable_io.py |
— |
| Table predicate / scan | Yes | predicate_filter (dense ~50% keep), predicate_filter_selective (~5–7%), scan_count |
Both keep-rate regimes; fused gather ≠ project+mask |
| Table schemas | Partial | mixed / narrow / wide / varlen | typed / ascii at selected row counts |
| Table GPU vs CPU | Partial | GPU transports / fitstable | Expanding into published tables |
| Writes / write→compress | Partial | suite expansion | Read-heavy historically; write parity also in tests |
| ML DataLoader | Yes | bench_ml_loader.py |
Reported in highlights / dedicated section |
Why the I/O transport table looks sparse on GPU¶
disk→GPUis always empty — backends decode on the host first, then.to(device). See Disk to GPU.disk→CPU→GPUvsdisk→RAM→GPU— mmap-off vs mmap-on host decode + H2D.- GPU rows need CUDA/MPS hardware — published CUDA numbers come from
CANFAR staging (
exhaustive_cuda_20260807_013736, refreshed 2026-08-07). - Tables — see Tables on GPU transports.
GPU integer dtype comparisons¶
The deficit table compares default
torchfits.read(..., scale_on_device=True) against
torch.from_numpy(fitsio.read(...)).to(cuda). That pairing is not
dtype-equivalent for every scaled integer FITS file.
| FITS convention | fitsio @ CUDA | default read @ CUDA |
|---|---|---|
| Signed byte (BITPIX=8, BZERO=-128) | native int8 H2D |
narrow int8 H2D + offset on device |
| Unsigned uint16/uint32 (BZERO) | native uint H2D | narrow storage H2D, offset on device |
| Generic BSCALE/BZERO | often native storage | float32 on device (ML-friendly) |
For apples-to-apples integer GPU timing, the suite also records
torchfits_dtype_fair_device (read_tensor(..., raw_scale=True)).
Training loops: call
torchfits.cache.optimize_for_dataset(paths, avg_file_size_mb=…) before
DataLoader epochs so handle caches stay warm.
Refreshing GPU numbers (CANFAR staging)¶
CUDA lab numbers come from a headless GPU session on @staging. From a
machine with canfar x509 auth:
bash scripts/selfcheck_canfar_launcher.sh
TORCHFITS_CANFAR_IMAGE=astroai/notebook:latest TORCHFITS_BENCH_MODE=exhaustive \
pixi run bench-canfar-gpu
bash scripts/fetch_canfar_bench_vos.sh exhaustive_cuda_<stamp>
bash scripts/patch_canfar_exhaustive_docs.sh exhaustive_cuda_<stamp>
# Local CI + docs before push
bash scripts/ci_local.sh
# Apple Silicon (MPS transport rows)
pixi run bench-mps
I/O transport and backend¶
GPU summary: Tensor
disk→CPU→GPU/disk→RAM→GPUrows appear only when the CSV was produced on CUDA or MPS.disk→GPUstays empty (unsupported). Table GPU cells stay empty until the table-GPU suite lands.
Source: docs/assets/bench/exhaustive_cpu_20260807_013736/results.csv (mmap on+off matrix.)
Cell values are median wall-clock over all comparable OK rows in the
(domain × I/O transport × backend) bucket; throughput is intentionally
omitted because the cell aggregates heterogeneous payloads and would
produce physically-impossible rates when small and large sizes are
median-mixed. See scripts/render_bench_iopath_table.py for the
aggregation rules.
Tensor I/O (IMAGE HDU) (fits)¶
| I/O transport | torchfits (libcfitsio) |
astropy |
fitsio |
cfitsio (direct) |
|---|---|---|---|---|
disk→CPU |
0.09 ms (n=174) |
0.59 ms (n=253) |
0.17 ms (n=261) |
— (engine exposed under torchfits) |
disk→RAM→CPU |
0.11 ms (n=174) |
0.46 ms (n=184) |
— (rows skipped under strict_mmap_fairness) |
— (engine exposed under torchfits) |
disk→GPU |
— | — | — | — |
disk→CPU→GPU |
— | — | — | — |
disk→RAM→GPU |
— | — | — | — |
Table I/O (fitstable)¶
| I/O transport | torchfits (libcfitsio) |
astropy |
fitsio |
cfitsio (direct) |
|---|---|---|---|---|
disk→CPU |
0.29 ms (n=216) |
3.30 ms (n=184) |
0.72 ms (n=216) |
— (engine exposed under torchfits) |
disk→RAM→CPU |
0.26 ms (n=216) |
3.25 ms (n=184) |
— (rows skipped under strict_mmap_fairness) |
— (engine exposed under torchfits) |
disk→GPU |
— | — | — | — |
disk→CPU→GPU |
— | — | — | — |
disk→RAM→GPU |
— | — | — | — |
Notes on the layout¶
- Rows are I/O transports (
disk→CPU,disk→RAM→CPU,disk→GPU,disk→CPU→GPU,disk→RAM→GPU). - Columns are backends (
torchfits/astropy/fitsio/cfitsio-direct). - Pure-C CFITSIO (vendored):
pixi run bench-cfitsio-directruns the full image+table benchmark fixture set with op→API mapping inbenchmarks/cfitsio_direct/bench_cfitsio_direct.c(fits_read_img/fits_read_subset/fits_read_record/fits_read_tblbytes/fits_read_col). CSV:benchmarks_results/<run-id>/cfitsio_direct.csv. - Cell
n=counts comparable OK rows in the bucket;—indicates the bucket is empty (no rows match, or rows were excluded understrict_mmap_fairnessin the originalbench-allsummary). - Median is computed over heterogeneous operations (
read_full,cutout_100x100,header_read,predicate_filter,projection,row_slice, etc.) and payload sizes; treat the per-cell ms as a coarse representative number, not a precise benchmark.
Performance highlights¶
The following table showcases median wall-clock times for key FITS tensor and table cases. The specialized column is torchfits_specialized (open-once / subset-reader paths); it is empty when that path was not measured.
| Benchmark Case | Device | torchfits | torchfits (specialized) | astropy (via torch) | fitsio (via torch) | Win vs Astropy | Win vs fitsio |
|---|---|---|---|---|---|---|---|
| Table read (100k rows, 8 cols, mixed) | CPU | 3.38 ms | 3.30 ms | 52.55 ms | 16.85 ms | 15.93x | 5.11x |
| Varlen table read (100k rows, 3 cols) | CPU | 74.74 ms | 10.80 ms | 528.67 ms | 107.37 ms | 48.93x | 9.94x |
Benchmark category summary¶
CPU category rows aggregate the CPU exhaustive
(exhaustive_cpu_20260807_013736, source of the generated
highlights and full table
above); the GPU (CUDA) rows come from exhaustive_cuda_20260807_013736
(see Published runs by platform below — all lags are listed, floors
labeled as noise vs significant). Category ranges are the last regenerated aggregation
shape; for absolute times prefer the generated tables above.
FITS image I/O¶
| Category | Cases | torchfits median | astropy median | fitsio median | Typical speedup vs astropy | Typical speedup vs fitsio |
|---|---|---|---|---|---|---|
| 1D (float32/64, int8–int64, tiny–large) | 24 | 29 μs – 1.28 ms | 302 μs – 2.57 ms | 61 μs – 1.69 ms | 2.0–13.5× | 1.30–2.4× |
| 2D (float32/64, int8–int64, uint16/32, tiny–large) | 30 | 37 μs – 7.07 ms | 361 μs – 13.10 ms | 75 μs – 8.92 ms | 1.8–12.9× | 1.26–2.2× |
| 3D (float32/64, int8–int64, tiny–medium) | 18 | 45 μs – 1.96 ms | 423 μs – 4.35 ms | 83 μs – 2.98 ms | 2.2–15.4× | 1.41–2.1× |
| Compressed (gzip, hcompress, rice) | 5 | 1.33–45.56 ms | 11.43–72.27 ms | 1.38–44.34 ms | 1.1–8.6× | 0.58–1.1× |
| Scaled (BSCALE/BZERO, small–large) | 3 | 76 μs – 2.93 ms | 496 μs – 6.09 ms | 128 μs – 3.87 ms | 2.1–6.6× | 1.32–1.7× |
| MEF (multi-extension, small/medium) | 2 | 69–220 μs | 614–954 μs | 155–342 μs | 4.3–8.9× | 1.55–2.2× |
| Multi-MEF (10 extensions, cutouts + random reads) | 3 | 64 μs – 8.15 ms | 634 μs – 12.01 ms | 194 μs – 11.04 ms | 1.5–40.4× | 1.36–3.0× |
| Repeated cutouts (50× 100×100) | 1 | 698 μs | 88.49 ms | 5.33 ms | 126.7× | 7.63× |
| Time series frames (5 frames) | 5 | 64–91 μs | 492–663 μs | 143–195 μs | 6.5–7.8× | 1.95–2.3× |
| Header read (all fixture types) | 87 | 14–51 μs | 257 μs – 2.46 ms | 25–267 μs | 15.1–48.2× | 1.58–5.2× |
GPU (CUDA) results — 85 comparable read_full / cutout cases:
| Category | torchfits median | astropy median | fitsio median | Typical speedup vs astropy | Typical speedup vs fitsio |
|---|---|---|---|---|---|
| 1D (tiny–large) | 107 μs – 1.57 ms | 438 μs – 2.90 ms | 106 μs – 2.05 ms | 1.8–6.8× | 0.88–1.5× |
| 2D (tiny–large) | 114 μs – 11.92 ms | 450 μs – 22.88 ms | 112 μs – 13.09 ms | 1.9–6.9× | 0.91–1.5× |
| 3D (tiny–medium) | 111 μs – 2.76 ms | 479 μs – 5.75 ms | 110 μs – 3.55 ms | 1.8–7.4× | 0.95–1.6× |
| Compressed (gzip, hcompress, rice) | 989 μs – 30.66 ms | 9.63–67.38 ms | 1.03–29.60 ms | 1.2–9.7× | 0.97–1.1× |
| Scaled | 198 μs – 4.34 ms | 920 μs – 10.93 ms | 203 μs – 4.91 ms | 1.9–4.6× | 1.03–1.1× |
| MEF + Multi-MEF | 143–391 μs | 1.14–2.80 ms | 173–451 μs | 4.0–17.0× | 1.16–1.7× |
| Repeated cutouts (GPU) | 1.47 ms | 61.22 ms | 6.19 ms | 41.7× | 4.21× |
FITS table I/O¶
| Category | Cases | torchfits median | astropy median | fitsio median | Typical speedup vs astropy | Typical speedup vs fitsio |
|---|---|---|---|---|---|---|
| read_full (all schemas incl. 1M rows, varlen) | 18 | 175 μs – 74.74 ms | 2.40–681.18 ms | 0.26–485.33 ms | 1.3–32× | 0.73–22× |
| projection (column subset) | 18 | 171 μs – 72.73 ms | 2.29–526.19 ms | 0.30–107.85 ms | 1.4–38× | 1.48–13× |
| row_slice (row range) | 18 | 113 μs – 6.96 ms | 1.99–59.12 ms | 0.25–31.15 ms | 7.8–55× | 1.55–17× |
| predicate_filter (WHERE clause, dense + selective) | 36 | 71 μs – 11.13 ms | 1.51–2.64 ms | 0.12–20.44 ms | 6.9–11× | 0.98–3× |
| scan_count (streaming) | 18 | 26–65 μs | 0.40–0.98 ms | 0.06–0.43 ms | 14.4–17× | 2.12–7× |
Exhaustive Benchmark Results¶
The complete, un-cherrypicked list of all measured configurations. Empty cells mean that method was not run for the case (for example torchfits_specialized is only used for open-once / subset-reader paths). Domain tensor = IMAGE HDU payloads (1D–4D); table = binary/ASCII tables.
| Domain | Benchmark Case | Operation | Size | Device | mmap | torchfits | torchfits (specialized) | astropy (via torch) | fitsio (via torch) | cfitsio (direct) | Speedup vs Astropy | Speedup vs fitsio |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| tensor | compressed_gzip_1:header_read | header_read | 1.29 MB | CPU | n/a | — | 50.8 μs | 2.34 ms | 237.1 μs | — | 46.08x | 4.66x |
| tensor | compressed_gzip_2:header_read | header_read | 0.89 MB | CPU | n/a | — | 49.3 μs | 2.36 ms | 233.4 μs | — | 47.97x | 4.74x |
| tensor | compressed_hcompress_1:header_read | header_read | 0.82 MB | CPU | n/a | — | 51.1 μs | 2.46 ms | 267.0 μs | — | 48.20x | 5.22x |
| tensor | compressed_rice_1:cutout_100x100 | cutout_100x100 | 0.90 MB | CPU | n/a | 1.33 ms | 1.27 ms | 11.55 ms | 1.37 ms | — | 9.08x | 1.08x |
| tensor | compressed_rice_1:header_read | header_read | 0.90 MB | CPU | n/a | — | 37.9 μs | 1.77 ms | 192.8 μs | — | 46.68x | 5.09x |
| tensor | large_float32_1d:header_read | header_read | 3.82 MB | CPU | n/a | — | 16.8 μs | 272.8 μs | 30.1 μs | — | 16.23x | 1.79x |
| tensor | large_float32_2d:header_read | header_read | 16.00 MB | CPU | n/a | — | 16.9 μs | 299.1 μs | 30.4 μs | — | 17.69x | 1.80x |
| tensor | large_float64_1d:header_read | header_read | 7.63 MB | CPU | n/a | — | 15.1 μs | 273.8 μs | 26.7 μs | — | 18.08x | 1.76x |
| tensor | large_float64_2d:header_read | header_read | 32.00 MB | CPU | n/a | — | 17.5 μs | 301.7 μs | 30.3 μs | — | 17.28x | 1.73x |
| tensor | large_int16_1d:header_read | header_read | 1.91 MB | CPU | n/a | — | 16.3 μs | 266.6 μs | 27.8 μs | — | 16.32x | 1.70x |
| tensor | large_int16_2d:header_read | header_read | 8.00 MB | CPU | n/a | — | 16.5 μs | 297.0 μs | 30.5 μs | — | 18.04x | 1.85x |
| tensor | large_int32_1d:header_read | header_read | 3.82 MB | CPU | n/a | — | 16.8 μs | 266.2 μs | 27.2 μs | — | 15.86x | 1.62x |
| tensor | large_int32_2d:header_read | header_read | 16.00 MB | CPU | n/a | — | 15.6 μs | 295.3 μs | 28.2 μs | — | 18.91x | 1.81x |
| tensor | large_int64_1d:header_read | header_read | 7.63 MB | CPU | n/a | — | 14.9 μs | 271.0 μs | 28.3 μs | — | 18.13x | 1.90x |
| tensor | large_int64_2d:header_read | header_read | 32.00 MB | CPU | n/a | — | 17.1 μs | 286.1 μs | 28.3 μs | — | 16.75x | 1.66x |
| tensor | large_int8_1d:header_read | header_read | 0.96 MB | CPU | n/a | — | 16.4 μs | 317.2 μs | 33.0 μs | — | 19.40x | 2.02x |
| tensor | large_int8_2d:header_read | header_read | 4.00 MB | CPU | n/a | — | 17.2 μs | 339.1 μs | 32.9 μs | — | 19.72x | 1.91x |
| tensor | large_uint16_2d:header_read | header_read | 8.00 MB | CPU | n/a | — | 17.5 μs | 339.0 μs | 35.4 μs | — | 19.36x | 2.02x |
| tensor | large_uint32_2d:header_read | header_read | 16.00 MB | CPU | n/a | — | 18.6 μs | 347.8 μs | 42.0 μs | — | 18.74x | 2.26x |
| tensor | medium_float32_1d:header_read | header_read | 0.38 MB | CPU | n/a | — | 17.9 μs | 269.2 μs | 29.9 μs | — | 15.05x | 1.67x |
| tensor | medium_float32_2d:header_read | header_read | 4.00 MB | CPU | n/a | — | 16.6 μs | 294.1 μs | 29.4 μs | — | 17.67x | 1.77x |
| tensor | medium_float32_3d:header_read | header_read | 6.25 MB | CPU | n/a | — | 18.2 μs | 320.0 μs | 31.4 μs | — | 17.57x | 1.72x |
| tensor | medium_float64_1d:header_read | header_read | 0.77 MB | CPU | n/a | — | 16.8 μs | 282.8 μs | 28.9 μs | — | 16.82x | 1.72x |
| tensor | medium_float64_2d:header_read | header_read | 8.00 MB | CPU | n/a | — | 18.2 μs | 303.4 μs | 28.7 μs | — | 16.71x | 1.58x |
| tensor | medium_float64_3d:header_read | header_read | 12.51 MB | CPU | n/a | — | 16.1 μs | 315.5 μs | 31.1 μs | — | 19.56x | 1.93x |
| tensor | medium_int16_1d:header_read | header_read | 0.20 MB | CPU | n/a | — | 15.7 μs | 269.5 μs | 27.8 μs | — | 17.16x | 1.77x |
| tensor | medium_int16_2d:header_read | header_read | 2.01 MB | CPU | n/a | — | 18.0 μs | 295.8 μs | 32.2 μs | — | 16.43x | 1.79x |
| tensor | medium_int16_3d:header_read | header_read | 3.13 MB | CPU | n/a | — | 17.8 μs | 317.1 μs | 31.8 μs | — | 17.85x | 1.79x |
| tensor | medium_int32_1d:header_read | header_read | 0.38 MB | CPU | n/a | — | 14.9 μs | 273.2 μs | 28.0 μs | — | 18.28x | 1.87x |
| tensor | medium_int32_2d:header_read | header_read | 4.00 MB | CPU | n/a | — | 16.3 μs | 294.6 μs | 29.0 μs | — | 18.03x | 1.77x |
| tensor | medium_int32_3d:header_read | header_read | 6.25 MB | CPU | n/a | — | 15.5 μs | 312.1 μs | 29.4 μs | — | 20.10x | 1.90x |
| tensor | medium_int64_1d:header_read | header_read | 0.77 MB | CPU | n/a | — | 16.6 μs | 278.7 μs | 26.5 μs | — | 16.75x | 1.59x |
| tensor | medium_int64_2d:header_read | header_read | 8.00 MB | CPU | n/a | — | 17.0 μs | 300.8 μs | 29.9 μs | — | 17.66x | 1.75x |
| tensor | medium_int64_3d:header_read | header_read | 12.51 MB | CPU | n/a | — | 15.6 μs | 314.6 μs | 32.3 μs | — | 20.19x | 2.08x |
| tensor | medium_int8_1d:header_read | header_read | 0.10 MB | CPU | n/a | — | 16.9 μs | 309.6 μs | 33.1 μs | — | 18.35x | 1.96x |
| tensor | medium_int8_2d:header_read | header_read | 1.01 MB | CPU | n/a | — | 16.9 μs | 330.5 μs | 35.3 μs | — | 19.55x | 2.09x |
| tensor | medium_int8_3d:header_read | header_read | 1.57 MB | CPU | n/a | — | 16.7 μs | 347.4 μs | 36.9 μs | — | 20.76x | 2.21x |
| tensor | medium_uint16_2d:header_read | header_read | 2.01 MB | CPU | n/a | — | 16.3 μs | 334.2 μs | 34.7 μs | — | 20.54x | 2.13x |
| tensor | medium_uint32_2d:header_read | header_read | 4.00 MB | CPU | n/a | — | 17.0 μs | 329.8 μs | 37.5 μs | — | 19.45x | 2.21x |
| tensor | mef_medium:header_read | header_read | 7.02 MB | CPU | n/a | — | 19.3 μs | 542.0 μs | 42.4 μs | — | 28.13x | 2.20x |
| tensor | mef_small:header_read | header_read | 0.45 MB | CPU | n/a | — | 18.3 μs | 528.1 μs | 42.7 μs | — | 28.90x | 2.34x |
| tensor | multi_mef_10ext:cutout_100x100 | cutout_100x100 | 2.68 MB | CPU | n/a | 96.4 μs | 129.7 μs | 3.79 ms | 275.6 μs | — | 39.32x | 2.86x |
| tensor | multi_mef_10ext:header_read | header_read | 2.68 MB | CPU | n/a | — | 19.7 μs | 528.6 μs | 42.8 μs | — | 26.87x | 2.18x |
| tensor | multi_mef_10ext:random_ext_full_reads_200 | random_ext_full_reads_200 | 2.68 MB | CPU | n/a | 8.15 ms | 8.15 ms | 11.89 ms | 11.03 ms | — | 1.46x | 1.35x |
| tensor | repeated_cutouts_50x_100x100:repeated_cutouts_50x_100x100 | repeated_cutouts_50x_100x100 | 4.00 MB | CPU | n/a | 698.5 μs | 682.3 μs | 88.41 ms | 5.13 ms | — | 129.59x | 7.52x |
| tensor | scaled_large:header_read | header_read | 8.00 MB | CPU | n/a | — | 29.6 μs | 642.7 μs | 66.4 μs | — | 21.69x | 2.24x |
| tensor | scaled_medium:header_read | header_read | 2.01 MB | CPU | n/a | — | 20.3 μs | 394.2 μs | 41.9 μs | — | 19.46x | 2.07x |
| tensor | scaled_small:header_read | header_read | 0.13 MB | CPU | n/a | — | 16.4 μs | 332.8 μs | 34.6 μs | — | 20.34x | 2.12x |
| tensor | small_float32_1d:header_read | header_read | 42.2 KB | CPU | n/a | — | 16.1 μs | 272.6 μs | 28.2 μs | — | 16.89x | 1.75x |
| tensor | small_float32_2d:header_read | header_read | 0.26 MB | CPU | n/a | — | 14.7 μs | 290.4 μs | 30.7 μs | — | 19.71x | 2.08x |
| tensor | small_float32_3d:header_read | header_read | 0.63 MB | CPU | n/a | — | 17.2 μs | 310.5 μs | 30.8 μs | — | 18.01x | 1.79x |
| tensor | small_float64_1d:header_read | header_read | 0.08 MB | CPU | n/a | — | 15.1 μs | 265.9 μs | 27.5 μs | — | 17.59x | 1.82x |
| tensor | small_float64_2d:header_read | header_read | 0.51 MB | CPU | n/a | — | 15.2 μs | 291.1 μs | 29.8 μs | — | 19.20x | 1.97x |
| tensor | small_float64_3d:header_read | header_read | 1.26 MB | CPU | n/a | — | 16.6 μs | 315.1 μs | 31.3 μs | — | 19.03x | 1.89x |
| tensor | small_int16_1d:header_read | header_read | 22.5 KB | CPU | n/a | — | 15.3 μs | 268.7 μs | 25.8 μs | — | 17.55x | 1.69x |
| tensor | small_int16_2d:header_read | header_read | 0.13 MB | CPU | n/a | — | 14.3 μs | 283.6 μs | 26.8 μs | — | 19.87x | 1.88x |
| tensor | small_int16_3d:header_read | header_read | 0.32 MB | CPU | n/a | — | 16.1 μs | 308.3 μs | 30.5 μs | — | 19.14x | 1.89x |
| tensor | small_int32_1d:header_read | header_read | 42.2 KB | CPU | n/a | — | 15.8 μs | 257.1 μs | 26.3 μs | — | 16.27x | 1.67x |
| tensor | small_int32_2d:header_read | header_read | 0.26 MB | CPU | n/a | — | 15.8 μs | 289.1 μs | 27.8 μs | — | 18.24x | 1.75x |
| tensor | small_int32_3d:header_read | header_read | 0.63 MB | CPU | n/a | — | 16.8 μs | 298.1 μs | 29.7 μs | — | 17.78x | 1.77x |
| tensor | small_int64_1d:header_read | header_read | 0.08 MB | CPU | n/a | — | 14.9 μs | 266.3 μs | 27.2 μs | — | 17.81x | 1.82x |
| tensor | small_int64_2d:header_read | header_read | 0.51 MB | CPU | n/a | — | 14.7 μs | 285.4 μs | 28.6 μs | — | 19.36x | 1.94x |
| tensor | small_int64_3d:header_read | header_read | 1.26 MB | CPU | n/a | — | 15.2 μs | 312.8 μs | 30.8 μs | — | 20.61x | 2.03x |
| tensor | small_int8_1d:header_read | header_read | 14.1 KB | CPU | n/a | — | 16.9 μs | 314.8 μs | 32.7 μs | — | 18.66x | 1.94x |
| tensor | small_int8_2d:header_read | header_read | 0.07 MB | CPU | n/a | — | 17.7 μs | 341.0 μs | 34.0 μs | — | 19.29x | 1.92x |
| tensor | small_int8_3d:header_read | header_read | 0.16 MB | CPU | n/a | — | 17.1 μs | 356.7 μs | 35.0 μs | — | 20.91x | 2.05x |
| tensor | small_uint16_2d:header_read | header_read | 0.13 MB | CPU | n/a | — | 16.2 μs | 337.0 μs | 33.0 μs | — | 20.76x | 2.04x |
| tensor | small_uint32_2d:header_read | header_read | 0.26 MB | CPU | n/a | — | 16.3 μs | 344.9 μs | 34.0 μs | — | 21.21x | 2.09x |
| tensor | timeseries_frame_000:header_read | header_read | 0.26 MB | CPU | n/a | — | 17.1 μs | 295.3 μs | 31.5 μs | — | 17.29x | 1.84x |
| tensor | timeseries_frame_001:header_read | header_read | 0.26 MB | CPU | n/a | — | 16.9 μs | 297.6 μs | 31.2 μs | — | 17.62x | 1.85x |
| tensor | timeseries_frame_002:header_read | header_read | 0.26 MB | CPU | n/a | — | 16.8 μs | 297.4 μs | 30.7 μs | — | 17.68x | 1.83x |
| tensor | timeseries_frame_003:header_read | header_read | 0.26 MB | CPU | n/a | — | 15.9 μs | 295.3 μs | 30.3 μs | — | 18.58x | 1.91x |
| tensor | timeseries_frame_004:header_read | header_read | 0.26 MB | CPU | n/a | — | 19.4 μs | 333.5 μs | 31.5 μs | — | 17.20x | 1.62x |
| tensor | tiny_float32_1d:header_read | header_read | 8.4 KB | CPU | n/a | — | 15.7 μs | 278.2 μs | 26.7 μs | — | 17.77x | 1.71x |
| tensor | tiny_float32_2d:header_read | header_read | 19.7 KB | CPU | n/a | — | 17.2 μs | 289.7 μs | 30.5 μs | — | 16.87x | 1.78x |
| tensor | tiny_float32_3d:header_read | header_read | 25.3 KB | CPU | n/a | — | 17.2 μs | 315.6 μs | 32.4 μs | — | 18.30x | 1.88x |
| tensor | tiny_float64_1d:header_read | header_read | 11.2 KB | CPU | n/a | — | 15.5 μs | 273.3 μs | 27.3 μs | — | 17.63x | 1.76x |
| tensor | tiny_float64_2d:header_read | header_read | 36.6 KB | CPU | n/a | — | 16.7 μs | 291.2 μs | 29.8 μs | — | 17.41x | 1.78x |
| tensor | tiny_float64_3d:header_read | header_read | 45.0 KB | CPU | n/a | — | 17.4 μs | 317.4 μs | 33.4 μs | — | 18.20x | 1.91x |
| tensor | tiny_int16_1d:header_read | header_read | 5.6 KB | CPU | n/a | — | 15.2 μs | 270.6 μs | 25.3 μs | — | 17.85x | 1.67x |
| tensor | tiny_int16_2d:header_read | header_read | 11.2 KB | CPU | n/a | — | 14.2 μs | 293.9 μs | 30.3 μs | — | 20.74x | 2.14x |
| tensor | tiny_int16_3d:header_read | header_read | 14.1 KB | CPU | n/a | — | 17.1 μs | 308.9 μs | 32.3 μs | — | 18.07x | 1.89x |
| tensor | tiny_int32_1d:header_read | header_read | 8.4 KB | CPU | n/a | — | 15.4 μs | 267.7 μs | 27.3 μs | — | 17.39x | 1.77x |
| tensor | tiny_int32_2d:header_read | header_read | 19.7 KB | CPU | n/a | — | 15.1 μs | 291.9 μs | 28.2 μs | — | 19.38x | 1.87x |
| tensor | tiny_int32_3d:header_read | header_read | 25.3 KB | CPU | n/a | — | 15.6 μs | 314.6 μs | 34.4 μs | — | 20.20x | 2.21x |
| tensor | tiny_int64_1d:header_read | header_read | 11.2 KB | CPU | n/a | — | 15.7 μs | 269.5 μs | 27.8 μs | — | 17.17x | 1.77x |
| tensor | tiny_int64_2d:header_read | header_read | 36.6 KB | CPU | n/a | — | 14.7 μs | 292.6 μs | 29.8 μs | — | 19.91x | 2.03x |
| tensor | tiny_int64_3d:header_read | header_read | 45.0 KB | CPU | n/a | — | 16.4 μs | 311.7 μs | 30.3 μs | — | 18.99x | 1.84x |
| tensor | tiny_int8_1d:header_read | header_read | 5.6 KB | CPU | n/a | — | 16.4 μs | 306.2 μs | 30.3 μs | — | 18.69x | 1.85x |
| tensor | tiny_int8_2d:header_read | header_read | 8.4 KB | CPU | n/a | — | 16.0 μs | 334.7 μs | 35.1 μs | — | 20.86x | 2.19x |
| tensor | tiny_int8_3d:header_read | header_read | 8.4 KB | CPU | n/a | — | 18.0 μs | 349.8 μs | 38.0 μs | — | 19.48x | 2.12x |
| tensor | write_compress_hcompress_medium_float32_2d | write_compress | 4.00 MB | CPU | n/a | 48.04 ms | — | 58.35 ms | — | — | 1.21x | — |
| tensor | write_compress_rice_medium_float32_2d | write_compress | 4.00 MB | CPU | n/a | 37.73 ms | — | 68.78 ms | — | — | 1.82x | — |
| tensor | compressed_gzip_1:read_full | read_full | 1.29 MB | CPU | off | 23.63 ms | 23.69 ms | 45.57 ms | 26.29 ms | — | 1.93x | 1.11x |
| tensor | compressed_gzip_2:read_full | read_full | 0.89 MB | CPU | off | 20.18 ms | 20.25 ms | 72.27 ms | 22.98 ms | — | 3.58x | 1.14x |
| tensor | compressed_hcompress_1:read_full | read_full | 0.82 MB | CPU | off | 45.56 ms | 45.54 ms | 51.44 ms | 44.34 ms | — | 1.13x | 0.97x |
| tensor | compressed_rice_1:read_full | read_full | 0.90 MB | CPU | off | 12.45 ms | 12.45 ms | 19.04 ms | 7.26 ms | — | 1.53x | 0.58x |
| tensor | large_float32_1d:read_full | read_full | 3.82 MB | CPU | off | 459.9 μs | 463.5 μs | 1.11 ms | 736.3 μs | — | 2.41x | 1.60x |
| tensor | large_float32_2d:read_full | read_full | 16.00 MB | CPU | off | 2.57 ms | 2.51 ms | 9.33 ms | 3.09 ms | — | 3.71x | 1.23x |
| tensor | large_float64_1d:read_full | read_full | 7.63 MB | CPU | off | 796.8 μs | 1.30 ms | 1.83 ms | 1.16 ms | — | 2.29x | 1.46x |
| tensor | large_float64_2d:read_full | read_full | 32.00 MB | CPU | off | 5.59 ms | 5.30 ms | 9.84 ms | 4.93 ms | — | 1.86x | 0.93x |
| tensor | large_int16_1d:read_full | read_full | 1.91 MB | CPU | off | 257.0 μs | 349.3 μs | 719.3 μs | 333.0 μs | — | 2.80x | 1.30x |
| tensor | large_int16_2d:read_full | read_full | 8.00 MB | CPU | off | 914.1 μs | 931.4 μs | 3.64 ms | 1.19 ms | — | 3.99x | 1.30x |
| tensor | large_int32_1d:read_full | read_full | 3.82 MB | CPU | off | 530.6 μs | 431.1 μs | 1.08 ms | 711.4 μs | — | 2.51x | 1.65x |
| tensor | large_int32_2d:read_full | read_full | 16.00 MB | CPU | off | 1.61 ms | 2.38 ms | 9.22 ms | 2.91 ms | — | 5.74x | 1.81x |
| tensor | large_int64_1d:read_full | read_full | 7.63 MB | CPU | off | 1.23 ms | 811.0 μs | 1.83 ms | 1.16 ms | — | 2.26x | 1.43x |
| tensor | large_int64_2d:read_full | read_full | 32.00 MB | CPU | off | 4.73 ms | 4.53 ms | 10.23 ms | 4.94 ms | — | 2.26x | 1.09x |
| tensor | large_int8_1d:read_full | read_full | 0.96 MB | CPU | off | 159.8 μs | 166.0 μs | 624.9 μs | 172.4 μs | — | 3.91x | 1.08x |
| tensor | large_int8_2d:read_full | read_full | 4.00 MB | CPU | off | 505.7 μs | 532.7 μs | 1.42 ms | 636.2 μs | — | 2.82x | 1.26x |
| tensor | large_uint16_2d:read_full | read_full | 8.00 MB | CPU | off | 1.19 ms | 1.19 ms | 4.01 ms | 1.49 ms | — | 3.38x | 1.25x |
| tensor | large_uint32_2d:read_full | read_full | 16.00 MB | CPU | off | 2.91 ms | 3.10 ms | 6.53 ms | 3.49 ms | — | 2.24x | 1.20x |
| tensor | medium_float32_1d:read_full | read_full | 0.38 MB | CPU | off | 48.1 μs | 77.2 μs | 358.6 μs | 107.3 μs | — | 7.46x | 2.23x |
| tensor | medium_float32_2d:read_full | read_full | 4.00 MB | CPU | off | 492.9 μs | 504.6 μs | 1.20 ms | 797.3 μs | — | 2.43x | 1.62x |
| tensor | medium_float32_3d:read_full | read_full | 6.25 MB | CPU | off | 679.6 μs | 866.5 μs | 1.61 ms | 1.12 ms | — | 2.37x | 1.65x |
| tensor | medium_float64_1d:read_full | read_full | 0.77 MB | CPU | off | 87.3 μs | 123.7 μs | 465.9 μs | 148.5 μs | — | 5.34x | 1.70x |
| tensor | medium_float64_2d:read_full | read_full | 8.00 MB | CPU | off | 863.0 μs | 844.2 μs | 2.47 ms | 1.22 ms | — | 2.92x | 1.45x |
| tensor | medium_float64_3d:read_full | read_full | 12.51 MB | CPU | off | 1.27 ms | 1.84 ms | 4.12 ms | 2.16 ms | — | 3.23x | 1.70x |
| tensor | medium_int16_1d:read_full | read_full | 0.20 MB | CPU | off | 36.7 μs | 63.7 μs | 303.8 μs | 66.1 μs | — | 8.29x | 1.80x |
| tensor | medium_int16_2d:read_full | read_full | 2.01 MB | CPU | off | 288.2 μs | 333.1 μs | 776.9 μs | 370.8 μs | — | 2.70x | 1.29x |
| tensor | medium_int16_3d:read_full | read_full | 3.13 MB | CPU | off | 426.1 μs | 490.7 μs | 1.01 ms | 541.2 μs | — | 2.38x | 1.27x |
| tensor | medium_int32_1d:read_full | read_full | 0.38 MB | CPU | off | 83.4 μs | 78.7 μs | 357.2 μs | 104.6 μs | — | 4.54x | 1.33x |
| tensor | medium_int32_2d:read_full | read_full | 4.00 MB | CPU | off | 471.6 μs | 550.0 μs | 1.14 ms | 744.2 μs | — | 2.41x | 1.58x |
| tensor | medium_int32_3d:read_full | read_full | 6.25 MB | CPU | off | 730.4 μs | 690.0 μs | 1.63 ms | 1.15 ms | — | 2.36x | 1.66x |
| tensor | medium_int64_1d:read_full | read_full | 0.77 MB | CPU | off | 120.5 μs | 123.2 μs | 440.1 μs | 143.7 μs | — | 3.65x | 1.19x |
| tensor | medium_int64_2d:read_full | read_full | 8.00 MB | CPU | off | 849.7 μs | 947.7 μs | 2.48 ms | 1.23 ms | — | 2.92x | 1.45x |
| tensor | medium_int64_3d:read_full | read_full | 12.51 MB | CPU | off | 2.00 ms | 1.84 ms | 4.19 ms | 2.22 ms | — | 2.28x | 1.21x |
| tensor | medium_int8_1d:read_full | read_full | 0.10 MB | CPU | off | 53.1 μs | 53.1 μs | 383.7 μs | 56.6 μs | — | 7.23x | 1.07x |
| tensor | medium_int8_2d:read_full | read_full | 1.01 MB | CPU | off | 165.3 μs | 174.5 μs | 656.6 μs | 173.8 μs | — | 3.97x | 1.05x |
| tensor | medium_int8_3d:read_full | read_full | 1.57 MB | CPU | off | 252.0 μs | 255.3 μs | 842.0 μs | 295.5 μs | — | 3.34x | 1.17x |
| tensor | medium_uint16_2d:read_full | read_full | 2.01 MB | CPU | off | 386.7 μs | 348.0 μs | 1.30 ms | 426.1 μs | — | 3.75x | 1.22x |
| tensor | medium_uint32_2d:read_full | read_full | 4.00 MB | CPU | off | 623.7 μs | 619.4 μs | 1.73 ms | 913.2 μs | — | 2.79x | 1.47x |
| tensor | mef_medium:read_full | read_full | 7.02 MB | CPU | off | 166.9 μs | 192.1 μs | 849.7 μs | 209.9 μs | — | 5.09x | 1.26x |
| tensor | mef_small:read_full | read_full | 0.45 MB | CPU | off | 42.6 μs | 55.7 μs | 559.2 μs | 84.9 μs | — | 13.14x | 2.00x |
| tensor | multi_mef_10ext:read_full | read_full | 2.68 MB | CPU | off | 56.0 μs | 33.1 μs | 586.2 μs | 138.0 μs | — | 17.71x | 4.17x |
| tensor | scaled_large:read_full | read_full | 8.00 MB | CPU | off | 3.43 ms | 3.67 ms | 5.30 ms | 3.26 ms | — | 1.54x | 0.95x |
| tensor | scaled_medium:read_full | read_full | 2.01 MB | CPU | off | 626.6 μs | 643.8 μs | 1.35 ms | 770.4 μs | — | 2.15x | 1.23x |
| tensor | scaled_small:read_full | read_full | 0.13 MB | CPU | off | 63.0 μs | 83.6 μs | 461.0 μs | 91.5 μs | — | 7.32x | 1.45x |
| tensor | small_float32_1d:read_full | read_full | 42.2 KB | CPU | off | 33.0 μs | 39.7 μs | 275.1 μs | 47.6 μs | — | 8.35x | 1.44x |
| tensor | small_float32_2d:read_full | read_full | 0.26 MB | CPU | off | 69.1 μs | 55.3 μs | 357.7 μs | 87.0 μs | — | 6.47x | 1.57x |
| tensor | small_float32_3d:read_full | read_full | 0.63 MB | CPU | off | 79.6 μs | 109.2 μs | 440.9 μs | 147.8 μs | — | 5.54x | 1.86x |
| tensor | small_float64_1d:read_full | read_full | 0.08 MB | CPU | off | 45.8 μs | 33.9 μs | 285.9 μs | 49.8 μs | — | 8.44x | 1.47x |
| tensor | small_float64_2d:read_full | read_full | 0.51 MB | CPU | off | 67.2 μs | 95.6 μs | 396.5 μs | 109.9 μs | — | 5.90x | 1.63x |
| tensor | small_float64_3d:read_full | read_full | 1.26 MB | CPU | off | 181.7 μs | 185.1 μs | 604.3 μs | 236.1 μs | — | 3.33x | 1.30x |
| tensor | small_int16_1d:read_full | read_full | 22.5 KB | CPU | off | 30.0 μs | 39.8 μs | 257.2 μs | 40.9 μs | — | 8.58x | 1.36x |
| tensor | small_int16_2d:read_full | read_full | 0.13 MB | CPU | off | 39.2 μs | 56.8 μs | 306.6 μs | 54.7 μs | — | 7.82x | 1.40x |
| tensor | small_int16_3d:read_full | read_full | 0.32 MB | CPU | off | 50.3 μs | 81.9 μs | 374.2 μs | 86.4 μs | — | 7.44x | 1.72x |
| tensor | small_int32_1d:read_full | read_full | 42.2 KB | CPU | off | 28.0 μs | 40.8 μs | 266.2 μs | 48.5 μs | — | 9.50x | 1.73x |
| tensor | small_int32_2d:read_full | read_full | 0.26 MB | CPU | off | 63.8 μs | 67.6 μs | 347.7 μs | 82.6 μs | — | 5.45x | 1.29x |
| tensor | small_int32_3d:read_full | read_full | 0.63 MB | CPU | off | 67.9 μs | 98.6 μs | 444.0 μs | 150.3 μs | — | 6.54x | 2.21x |
| tensor | small_int64_1d:read_full | read_full | 0.08 MB | CPU | off | 43.0 μs | 42.9 μs | 278.8 μs | 48.6 μs | — | 6.50x | 1.13x |
| tensor | small_int64_2d:read_full | read_full | 0.51 MB | CPU | off | 95.9 μs | 78.8 μs | 396.8 μs | 109.8 μs | — | 5.03x | 1.39x |
| tensor | small_int64_3d:read_full | read_full | 1.26 MB | CPU | off | 187.4 μs | 159.3 μs | 603.4 μs | 232.7 μs | — | 3.79x | 1.46x |
| tensor | small_int8_1d:read_full | read_full | 14.1 KB | CPU | off | 28.4 μs | 48.4 μs | 361.6 μs | 42.5 μs | — | 12.74x | 1.50x |
| tensor | small_int8_2d:read_full | read_full | 0.07 MB | CPU | off | 44.5 μs | 38.4 μs | 391.9 μs | 55.0 μs | — | 10.20x | 1.43x |
| tensor | small_int8_3d:read_full | read_full | 0.16 MB | CPU | off | 59.6 μs | 60.8 μs | 426.7 μs | 65.5 μs | — | 7.16x | 1.10x |
| tensor | small_uint16_2d:read_full | read_full | 0.13 MB | CPU | off | 55.2 μs | 43.2 μs | 380.1 μs | 60.0 μs | — | 8.80x | 1.39x |
| tensor | small_uint32_2d:read_full | read_full | 0.26 MB | CPU | off | 41.9 μs | 74.0 μs | 432.1 μs | 94.2 μs | — | 10.32x | 2.25x |
| tensor | timeseries_frame_000:read_full | read_full | 0.26 MB | CPU | off | 126.9 μs | 135.4 μs | 562.9 μs | 137.4 μs | — | 4.44x | 1.08x |
| tensor | timeseries_frame_001:read_full | read_full | 0.26 MB | CPU | off | 76.2 μs | 90.3 μs | 562.7 μs | 139.6 μs | — | 7.39x | 1.83x |
| tensor | timeseries_frame_002:read_full | read_full | 0.26 MB | CPU | off | 91.1 μs | 84.8 μs | 562.0 μs | 136.5 μs | — | 6.62x | 1.61x |
| tensor | timeseries_frame_003:read_full | read_full | 0.26 MB | CPU | off | 89.8 μs | 71.3 μs | 553.5 μs | 129.7 μs | — | 7.77x | 1.82x |
| tensor | timeseries_frame_004:read_full | read_full | 0.26 MB | CPU | off | 82.6 μs | 90.1 μs | 559.9 μs | 136.3 μs | — | 6.78x | 1.65x |
| tensor | tiny_float32_1d:read_full | read_full | 8.4 KB | CPU | off | 65.0 μs | 51.7 μs | 438.9 μs | 72.4 μs | — | 8.48x | 1.40x |
| tensor | tiny_float32_2d:read_full | read_full | 19.7 KB | CPU | off | 58.6 μs | 62.9 μs | 471.5 μs | 70.7 μs | — | 8.05x | 1.21x |
| tensor | tiny_float32_3d:read_full | read_full | 25.3 KB | CPU | off | 48.2 μs | 57.1 μs | 496.0 μs | 75.6 μs | — | 10.30x | 1.57x |
| tensor | tiny_float64_1d:read_full | read_full | 11.2 KB | CPU | off | 59.7 μs | 65.5 μs | 456.7 μs | 71.3 μs | — | 7.65x | 1.19x |
| tensor | tiny_float64_2d:read_full | read_full | 36.6 KB | CPU | off | 51.2 μs | 65.3 μs | 488.5 μs | 79.6 μs | — | 9.54x | 1.55x |
| tensor | tiny_float64_3d:read_full | read_full | 45.0 KB | CPU | off | 61.2 μs | 67.2 μs | 508.5 μs | 77.3 μs | — | 8.31x | 1.26x |
| tensor | tiny_int16_1d:read_full | read_full | 5.6 KB | CPU | off | 64.2 μs | 46.6 μs | 424.3 μs | 69.4 μs | — | 9.11x | 1.49x |
| tensor | tiny_int16_2d:read_full | read_full | 11.2 KB | CPU | off | 50.4 μs | 62.8 μs | 463.0 μs | 74.6 μs | — | 9.19x | 1.48x |
| tensor | tiny_int16_3d:read_full | read_full | 14.1 KB | CPU | off | 66.0 μs | 64.2 μs | 500.2 μs | 74.6 μs | — | 7.79x | 1.16x |
| tensor | tiny_int32_1d:read_full | read_full | 8.4 KB | CPU | off | 60.2 μs | 46.6 μs | 440.9 μs | 71.8 μs | — | 9.47x | 1.54x |
| tensor | tiny_int32_2d:read_full | read_full | 19.7 KB | CPU | off | 44.1 μs | 51.6 μs | 482.5 μs | 77.3 μs | — | 10.93x | 1.75x |
| tensor | tiny_int32_3d:read_full | read_full | 25.3 KB | CPU | off | 53.7 μs | 61.6 μs | 509.3 μs | 78.4 μs | — | 9.49x | 1.46x |
| tensor | tiny_int64_1d:read_full | read_full | 11.2 KB | CPU | off | 58.9 μs | 52.5 μs | 455.9 μs | 75.9 μs | — | 8.68x | 1.44x |
| tensor | tiny_int64_2d:read_full | read_full | 36.6 KB | CPU | off | 64.2 μs | 45.7 μs | 470.8 μs | 70.3 μs | — | 10.30x | 1.54x |
| tensor | tiny_int64_3d:read_full | read_full | 45.0 KB | CPU | off | 47.6 μs | 69.0 μs | 496.1 μs | 78.2 μs | — | 10.43x | 1.64x |
| tensor | tiny_int8_1d:read_full | read_full | 5.6 KB | CPU | off | 49.4 μs | 40.8 μs | 611.3 μs | 78.1 μs | — | 14.98x | 1.91x |
| tensor | tiny_int8_2d:read_full | read_full | 8.4 KB | CPU | off | 55.9 μs | 53.6 μs | 638.6 μs | 79.6 μs | — | 11.92x | 1.49x |
| tensor | tiny_int8_3d:read_full | read_full | 8.4 KB | CPU | off | 45.0 μs | 64.6 μs | 659.0 μs | 80.1 μs | — | 14.64x | 1.78x |
| tensor | compressed_gzip_1:read_full | read_full | 1.29 MB | CPU | on | 23.77 ms | 23.70 ms | 46.38 ms | 26.41 ms | — | 1.96x | 1.11x |
| tensor | compressed_gzip_2:read_full | read_full | 0.89 MB | CPU | on | 20.32 ms | 20.36 ms | 72.52 ms | 23.06 ms | — | 3.57x | 1.13x |
| tensor | compressed_hcompress_1:read_full | read_full | 0.82 MB | CPU | on | 45.70 ms | 45.72 ms | 51.70 ms | 44.51 ms | — | 1.13x | 0.97x |
| tensor | compressed_rice_1:read_full | read_full | 0.90 MB | CPU | on | 12.65 ms | 12.63 ms | 33.26 ms | 12.64 ms | — | 2.63x | 1.00x |
| tensor | large_float32_1d:read_full | read_full | 3.82 MB | CPU | on | 667.0 μs | 803.9 μs | 1.49 ms | — | — | 2.24x | — |
| tensor | large_float32_2d:read_full | read_full | 16.00 MB | CPU | on | 3.91 ms | 4.03 ms | 13.50 ms | — | — | 3.45x | — |
| tensor | large_float64_1d:read_full | read_full | 7.63 MB | CPU | on | 1.28 ms | 1.23 ms | 2.46 ms | — | — | 2.00x | — |
| tensor | large_float64_2d:read_full | read_full | 32.00 MB | CPU | on | 8.56 ms | 7.74 ms | 12.67 ms | — | — | 1.64x | — |
| tensor | large_int16_1d:read_full | read_full | 1.91 MB | CPU | on | 373.6 μs | 398.4 μs | 1.04 ms | — | — | 2.78x | — |
| tensor | large_int16_2d:read_full | read_full | 8.00 MB | CPU | on | 1.37 ms | 1.33 ms | 2.58 ms | — | — | 1.93x | — |
| tensor | large_int32_1d:read_full | read_full | 3.82 MB | CPU | on | 907.4 μs | 848.8 μs | 1.48 ms | — | — | 1.74x | — |
| tensor | large_int32_2d:read_full | read_full | 16.00 MB | CPU | on | 4.08 ms | 3.95 ms | 12.84 ms | — | — | 3.25x | — |
| tensor | large_int64_1d:read_full | read_full | 7.63 MB | CPU | on | 1.32 ms | 1.35 ms | 2.40 ms | — | — | 1.82x | — |
| tensor | large_int64_2d:read_full | read_full | 32.00 MB | CPU | on | 6.93 ms | 6.76 ms | 11.45 ms | — | — | 1.69x | — |
| tensor | large_int8_1d:read_full | read_full | 0.96 MB | CPU | on | 250.4 μs | 253.6 μs | — | — | — | — | — |
| tensor | large_int8_2d:read_full | read_full | 4.00 MB | CPU | on | 828.3 μs | 786.8 μs | — | — | — | — | — |
| tensor | large_uint16_2d:read_full | read_full | 8.00 MB | CPU | on | 1.31 ms | 1.33 ms | — | — | — | — | — |
| tensor | large_uint32_2d:read_full | read_full | 16.00 MB | CPU | on | 5.33 ms | 4.53 ms | — | — | — | — | — |
| tensor | medium_float32_1d:read_full | read_full | 0.38 MB | CPU | on | 113.2 μs | 120.9 μs | 605.2 μs | — | — | 5.35x | — |
| tensor | medium_float32_2d:read_full | read_full | 4.00 MB | CPU | on | 709.9 μs | 724.4 μs | 1.57 ms | — | — | 2.22x | — |
| tensor | medium_float32_3d:read_full | read_full | 6.25 MB | CPU | on | 1.03 ms | 1.02 ms | 2.13 ms | — | — | 2.10x | — |
| tensor | medium_float64_1d:read_full | read_full | 0.77 MB | CPU | on | 174.5 μs | 174.7 μs | 705.2 μs | — | — | 4.04x | — |
| tensor | medium_float64_2d:read_full | read_full | 8.00 MB | CPU | on | 1.29 ms | 1.57 ms | 2.52 ms | — | — | 1.96x | — |
| tensor | medium_float64_3d:read_full | read_full | 12.51 MB | CPU | on | 1.93 ms | 2.46 ms | 3.68 ms | — | — | 1.91x | — |
| tensor | medium_int16_1d:read_full | read_full | 0.20 MB | CPU | on | 117.3 μs | 108.0 μs | 518.2 μs | — | — | 4.80x | — |
| tensor | medium_int16_2d:read_full | read_full | 2.01 MB | CPU | on | 428.1 μs | 388.4 μs | 1.07 ms | — | — | 2.75x | — |
| tensor | medium_int16_3d:read_full | read_full | 3.13 MB | CPU | on | 587.7 μs | 560.9 μs | 1.36 ms | — | — | 2.43x | — |
| tensor | medium_int32_1d:read_full | read_full | 0.38 MB | CPU | on | 103.9 μs | 147.4 μs | 587.3 μs | — | — | 5.65x | — |
| tensor | medium_int32_2d:read_full | read_full | 4.00 MB | CPU | on | 798.5 μs | 790.0 μs | 1.55 ms | — | — | 1.97x | — |
| tensor | medium_int32_3d:read_full | read_full | 6.25 MB | CPU | on | 1.29 ms | 1.21 ms | 2.15 ms | — | — | 1.77x | — |
| tensor | medium_int64_1d:read_full | read_full | 0.77 MB | CPU | on | 212.8 μs | 192.9 μs | 706.3 μs | — | — | 3.66x | — |
| tensor | medium_int64_2d:read_full | read_full | 8.00 MB | CPU | on | 1.56 ms | 1.31 ms | 2.52 ms | — | — | 1.92x | — |
| tensor | medium_int64_3d:read_full | read_full | 12.51 MB | CPU | on | 1.91 ms | 1.97 ms | 3.70 ms | — | — | 1.93x | — |
| tensor | medium_int8_1d:read_full | read_full | 0.10 MB | CPU | on | 68.7 μs | 92.6 μs | — | — | — | — | — |
| tensor | medium_int8_2d:read_full | read_full | 1.01 MB | CPU | on | 265.2 μs | 205.3 μs | — | — | — | — | — |
| tensor | medium_int8_3d:read_full | read_full | 1.57 MB | CPU | on | 371.2 μs | 397.3 μs | — | — | — | — | — |
| tensor | medium_uint16_2d:read_full | read_full | 2.01 MB | CPU | on | 419.4 μs | 414.2 μs | — | — | — | — | — |
| tensor | medium_uint32_2d:read_full | read_full | 4.00 MB | CPU | on | 1.18 ms | 1.18 ms | — | — | — | — | — |
| tensor | mef_medium:read_full | read_full | 7.02 MB | CPU | on | 273.6 μs | 259.8 μs | — | — | — | — | — |
| tensor | mef_small:read_full | read_full | 0.45 MB | CPU | on | 95.7 μs | 53.3 μs | — | — | — | — | — |
| tensor | multi_mef_10ext:read_full | read_full | 2.68 MB | CPU | on | 71.9 μs | 76.7 μs | — | — | — | — | — |
| tensor | scaled_large:read_full | read_full | 8.00 MB | CPU | on | 2.43 ms | 3.79 ms | — | — | — | — | — |
| tensor | scaled_medium:read_full | read_full | 2.01 MB | CPU | on | 650.3 μs | 650.9 μs | — | — | — | — | — |
| tensor | scaled_small:read_full | read_full | 0.13 MB | CPU | on | 88.2 μs | 88.2 μs | — | — | — | — | — |
| tensor | small_float32_1d:read_full | read_full | 42.2 KB | CPU | on | 25.2 μs | 41.4 μs | 285.5 μs | — | — | 11.32x | — |
| tensor | small_float32_2d:read_full | read_full | 0.26 MB | CPU | on | 61.8 μs | 45.9 μs | 366.7 μs | — | — | 7.99x | — |
| tensor | small_float32_3d:read_full | read_full | 0.63 MB | CPU | on | 102.9 μs | 113.1 μs | 456.6 μs | — | — | 4.44x | — |
| tensor | small_float64_1d:read_full | read_full | 0.08 MB | CPU | on | 30.9 μs | 44.4 μs | 285.0 μs | — | — | 9.24x | — |
| tensor | small_float64_2d:read_full | read_full | 0.51 MB | CPU | on | 86.0 μs | 74.3 μs | 412.2 μs | — | — | 5.55x | — |
| tensor | small_float64_3d:read_full | read_full | 1.26 MB | CPU | on | 176.0 μs | 179.3 μs | 580.3 μs | — | — | 3.30x | — |
| tensor | small_int16_1d:read_full | read_full | 22.5 KB | CPU | on | 37.8 μs | 49.5 μs | 275.5 μs | — | — | 7.29x | — |
| tensor | small_int16_2d:read_full | read_full | 0.13 MB | CPU | on | 55.4 μs | 78.2 μs | 321.3 μs | — | — | 5.80x | — |
| tensor | small_int16_3d:read_full | read_full | 0.32 MB | CPU | on | 87.8 μs | 84.4 μs | 385.4 μs | — | — | 4.57x | — |
| tensor | small_int32_1d:read_full | read_full | 42.2 KB | CPU | on | 42.3 μs | 51.5 μs | 279.9 μs | — | — | 6.62x | — |
| tensor | small_int32_2d:read_full | read_full | 0.26 MB | CPU | on | 88.0 μs | 75.0 μs | 355.8 μs | — | — | 4.75x | — |
| tensor | small_int32_3d:read_full | read_full | 0.63 MB | CPU | on | 119.2 μs | 112.5 μs | 458.2 μs | — | — | 4.07x | — |
| tensor | small_int64_1d:read_full | read_full | 0.08 MB | CPU | on | 41.8 μs | 62.5 μs | 315.3 μs | — | — | 7.54x | — |
| tensor | small_int64_2d:read_full | read_full | 0.51 MB | CPU | on | 102.7 μs | 74.1 μs | 422.8 μs | — | — | 5.70x | — |
| tensor | small_int64_3d:read_full | read_full | 1.26 MB | CPU | on | 195.7 μs | 171.6 μs | 580.9 μs | — | — | 3.38x | — |
| tensor | small_int8_1d:read_full | read_full | 14.1 KB | CPU | on | 30.2 μs | 50.7 μs | — | — | — | — | — |
| tensor | small_int8_2d:read_full | read_full | 0.07 MB | CPU | on | 54.3 μs | 30.1 μs | — | — | — | — | — |
| tensor | small_int8_3d:read_full | read_full | 0.16 MB | CPU | on | 68.7 μs | 34.5 μs | — | — | — | — | — |
| tensor | small_uint16_2d:read_full | read_full | 0.13 MB | CPU | on | 81.6 μs | 51.4 μs | — | — | — | — | — |
| tensor | small_uint32_2d:read_full | read_full | 0.26 MB | CPU | on | 78.8 μs | 110.3 μs | — | — | — | — | — |
| tensor | timeseries_frame_000:read_full | read_full | 0.26 MB | CPU | on | 55.8 μs | 49.9 μs | 364.3 μs | — | — | 7.30x | — |
| tensor | timeseries_frame_001:read_full | read_full | 0.26 MB | CPU | on | 66.7 μs | 61.2 μs | 354.8 μs | — | — | 5.80x | — |
| tensor | timeseries_frame_002:read_full | read_full | 0.26 MB | CPU | on | 36.2 μs | 66.3 μs | 361.0 μs | — | — | 9.98x | — |
| tensor | timeseries_frame_003:read_full | read_full | 0.26 MB | CPU | on | 60.8 μs | 63.1 μs | 361.4 μs | — | — | 5.95x | — |
| tensor | timeseries_frame_004:read_full | read_full | 0.26 MB | CPU | on | 64.0 μs | 43.7 μs | 355.9 μs | — | — | 8.14x | — |
| tensor | tiny_float32_1d:read_full | read_full | 8.4 KB | CPU | on | 27.6 μs | 27.7 μs | 275.2 μs | — | — | 9.97x | — |
| tensor | tiny_float32_2d:read_full | read_full | 19.7 KB | CPU | on | 29.4 μs | 31.8 μs | 294.8 μs | — | — | 10.02x | — |
| tensor | tiny_float32_3d:read_full | read_full | 25.3 KB | CPU | on | 41.9 μs | 25.1 μs | 300.1 μs | — | — | 11.97x | — |
| tensor | tiny_float64_1d:read_full | read_full | 11.2 KB | CPU | on | 22.9 μs | 39.6 μs | 273.6 μs | — | — | 11.97x | — |
| tensor | tiny_float64_2d:read_full | read_full | 36.6 KB | CPU | on | 33.1 μs | 26.9 μs | 287.0 μs | — | — | 10.67x | — |
| tensor | tiny_float64_3d:read_full | read_full | 45.0 KB | CPU | on | 32.5 μs | 35.4 μs | 306.4 μs | — | — | 9.44x | — |
| tensor | tiny_int16_1d:read_full | read_full | 5.6 KB | CPU | on | 31.6 μs | 32.4 μs | 265.2 μs | — | — | 8.40x | — |
| tensor | tiny_int16_2d:read_full | read_full | 11.2 KB | CPU | on | 49.5 μs | 26.4 μs | 287.2 μs | — | — | 10.90x | — |
| tensor | tiny_int16_3d:read_full | read_full | 14.1 KB | CPU | on | 44.1 μs | 39.0 μs | 291.8 μs | — | — | 7.48x | — |
| tensor | tiny_int32_1d:read_full | read_full | 8.4 KB | CPU | on | 47.4 μs | 36.5 μs | 272.7 μs | — | — | 7.48x | — |
| tensor | tiny_int32_2d:read_full | read_full | 19.7 KB | CPU | on | 30.5 μs | 36.0 μs | 292.6 μs | — | — | 9.61x | — |
| tensor | tiny_int32_3d:read_full | read_full | 25.3 KB | CPU | on | 46.5 μs | 39.4 μs | 302.1 μs | — | — | 7.67x | — |
| tensor | tiny_int64_1d:read_full | read_full | 11.2 KB | CPU | on | 33.3 μs | 29.1 μs | 268.3 μs | — | — | 9.22x | — |
| tensor | tiny_int64_2d:read_full | read_full | 36.6 KB | CPU | on | 30.4 μs | 53.7 μs | 290.1 μs | — | — | 9.55x | — |
| tensor | tiny_int64_3d:read_full | read_full | 45.0 KB | CPU | on | 46.6 μs | 49.2 μs | 306.9 μs | — | — | 6.59x | — |
| tensor | tiny_int8_1d:read_full | read_full | 5.6 KB | CPU | on | 46.1 μs | 36.3 μs | — | — | — | — | — |
| tensor | tiny_int8_2d:read_full | read_full | 8.4 KB | CPU | on | 51.6 μs | 29.3 μs | — | — | — | — | — |
| tensor | tiny_int8_3d:read_full | read_full | 8.4 KB | CPU | on | 46.2 μs | 29.4 μs | — | — | — | — | — |
| table | ascii_10000 | predicate_filter | 0.44 MB | CPU | off | 382.6 μs | 379.7 μs | 2.66 ms | 356.9 μs | — | 7.02x | 0.94x |
| table | ascii_10000 | predicate_filter_selective | 0.44 MB | CPU | off | 370.6 μs | 371.9 μs | 2.64 ms | 359.8 μs | — | 7.12x | 0.97x |
| table | ascii_10000 | projection | 0.44 MB | CPU | off | 1.02 ms | 984.1 μs | 7.98 ms | 1.97 ms | — | 8.11x | 2.00x |
| table | ascii_10000 | read_full | 0.44 MB | CPU | off | 1.01 ms | 979.2 μs | 8.00 ms | 1.96 ms | — | 8.17x | 2.00x |
| table | ascii_10000 | row_slice | 0.44 MB | CPU | off | 194.8 μs | 162.3 μs | 2.58 ms | 513.1 μs | — | 15.87x | 3.16x |
| table | ascii_10000 | scan_count | 0.44 MB | CPU | off | 25.6 μs | 24.4 μs | 403.5 μs | 66.6 μs | — | 16.57x | 2.73x |
| table | ascii_1000 | predicate_filter | 50.6 KB | CPU | off | 140.9 μs | 144.7 μs | 1.51 ms | 160.6 μs | — | 10.72x | 1.14x |
| table | ascii_1000 | predicate_filter_selective | 50.6 KB | CPU | off | 143.3 μs | 145.3 μs | 1.50 ms | 162.0 μs | — | 10.43x | 1.13x |
| table | ascii_1000 | projection | 50.6 KB | CPU | off | 188.5 μs | 167.5 μs | 2.15 ms | 333.5 μs | — | 12.85x | 1.99x |
| table | ascii_1000 | read_full | 50.6 KB | CPU | off | 184.1 μs | 170.5 μs | 2.16 ms | 321.9 μs | — | 12.65x | 1.89x |
| table | ascii_1000 | row_slice | 50.6 KB | CPU | off | 112.6 μs | 94.5 μs | 1.93 ms | 191.0 μs | — | 20.45x | 2.02x |
| table | ascii_1000 | scan_count | 50.6 KB | CPU | off | 26.7 μs | 27.5 μs | 415.4 μs | 63.0 μs | — | 15.53x | 2.36x |
| table | mixed_1000000 | predicate_filter | 50.55 MB | CPU | off | 11.13 ms | 10.52 ms | 15.06 ms | 20.26 ms | — | 1.43x | 1.93x |
| table | mixed_1000000 | predicate_filter_selective | 50.55 MB | CPU | off | 9.53 ms | 11.01 ms | 18.18 ms | 17.80 ms | — | 1.91x | 1.87x |
| table | mixed_1000000 | projection | 50.55 MB | CPU | off | 10.10 ms | 11.24 ms | 17.34 ms | 30.87 ms | — | 1.72x | 3.06x |
| table | mixed_1000000 | read_full | 50.55 MB | CPU | off | 30.32 ms | 30.54 ms | 334.99 ms | 116.95 ms | — | 11.05x | 3.86x |
| table | mixed_1000000 | row_slice | 50.55 MB | CPU | off | 290.8 μs | 286.1 μs | 12.78 ms | 1.57 ms | — | 44.67x | 5.48x |
| table | mixed_1000000 | scan_count | 50.55 MB | CPU | off | 38.3 μs | 30.8 μs | 447.2 μs | 80.5 μs | — | 14.52x | 2.61x |
| table | mixed_100000 | predicate_filter | 5.06 MB | CPU | off | 2.17 ms | 2.01 ms | 5.15 ms | 3.71 ms | — | 2.56x | 1.84x |
| table | mixed_100000 | predicate_filter_selective | 5.06 MB | CPU | off | 1.83 ms | 1.87 ms | 4.61 ms | 3.04 ms | — | 2.52x | 1.66x |
| table | mixed_100000 | projection | 5.06 MB | CPU | off | 1.79 ms | 1.75 ms | 5.13 ms | 5.35 ms | — | 2.92x | 3.05x |
| table | mixed_100000 | read_full | 5.06 MB | CPU | off | 3.38 ms | 3.30 ms | 52.55 ms | 16.85 ms | — | 15.93x | 5.11x |
| table | mixed_100000 | row_slice | 5.06 MB | CPU | off | 478.1 μs | 438.2 μs | 10.02 ms | 2.55 ms | — | 22.86x | 5.81x |
| table | mixed_100000 | scan_count | 5.06 MB | CPU | off | 45.2 μs | 48.2 μs | 767.5 μs | 134.3 μs | — | 16.99x | 2.97x |
| table | mixed_10000 | predicate_filter | 0.51 MB | CPU | off | 389.1 μs | 460.8 μs | 3.31 ms | 622.5 μs | — | 8.52x | 1.60x |
| table | mixed_10000 | predicate_filter_selective | 0.51 MB | CPU | off | 312.5 μs | 370.1 μs | 3.24 ms | 554.9 μs | — | 10.36x | 1.78x |
| table | mixed_10000 | projection | 0.51 MB | CPU | off | 292.1 μs | 246.3 μs | 3.28 ms | 782.3 μs | — | 13.30x | 3.18x |
| table | mixed_10000 | read_full | 0.51 MB | CPU | off | 466.7 μs | 440.5 μs | 7.84 ms | 1.83 ms | — | 17.80x | 4.14x |
| table | mixed_10000 | row_slice | 0.51 MB | CPU | off | 222.7 μs | 203.3 μs | 5.12 ms | 532.7 μs | — | 25.18x | 2.62x |
| table | mixed_10000 | scan_count | 0.51 MB | CPU | off | 28.1 μs | 29.6 μs | 440.7 μs | 76.0 μs | — | 15.71x | 2.71x |
| table | mixed_1000 | predicate_filter | 0.06 MB | CPU | off | 138.7 μs | 212.5 μs | 3.12 ms | 309.5 μs | — | 22.52x | 2.23x |
| table | mixed_1000 | predicate_filter_selective | 0.06 MB | CPU | off | 130.5 μs | 207.9 μs | 3.11 ms | 299.6 μs | — | 23.82x | 2.30x |
| table | mixed_1000 | projection | 0.06 MB | CPU | off | 176.4 μs | 143.2 μs | 3.13 ms | 322.1 μs | — | 21.83x | 2.25x |
| table | mixed_1000 | read_full | 0.06 MB | CPU | off | 217.5 μs | 195.1 μs | 3.73 ms | 442.8 μs | — | 19.10x | 2.27x |
| table | mixed_1000 | row_slice | 0.06 MB | CPU | off | 197.5 μs | 180.0 μs | 4.60 ms | 332.3 μs | — | 25.57x | 1.85x |
| table | mixed_1000 | scan_count | 0.06 MB | CPU | off | 46.5 μs | 46.7 μs | 784.1 μs | 123.3 μs | — | 16.85x | 2.65x |
| table | narrow_1000000 | predicate_filter | 12.40 MB | CPU | off | 6.27 ms | 5.58 ms | 8.48 ms | 14.73 ms | — | 1.52x | 2.64x |
| table | narrow_1000000 | predicate_filter_selective | 12.40 MB | CPU | off | 4.84 ms | 4.84 ms | 5.05 ms | 11.15 ms | — | 1.04x | 2.30x |
| table | narrow_1000000 | projection | 12.40 MB | CPU | off | 4.43 ms | 4.40 ms | 5.46 ms | 23.54 ms | — | 1.24x | 5.34x |
| table | narrow_1000000 | read_full | 12.40 MB | CPU | off | 6.00 ms | 5.93 ms | 6.69 ms | 5.85 ms | — | 1.13x | 0.99x |
| table | narrow_1000000 | row_slice | 12.40 MB | CPU | off | 159.8 μs | 148.5 μs | 3.84 ms | 566.0 μs | — | 25.87x | 3.81x |
| table | narrow_1000000 | scan_count | 12.40 MB | CPU | off | 28.1 μs | 29.1 μs | 440.9 μs | 67.5 μs | — | 15.69x | 2.40x |
| table | narrow_100000 | predicate_filter | 1.25 MB | CPU | off | 1.33 ms | 1.17 ms | 3.58 ms | 2.79 ms | — | 3.06x | 2.38x |
| table | narrow_100000 | predicate_filter_selective | 1.25 MB | CPU | off | 970.3 μs | 1.01 ms | 2.93 ms | 2.16 ms | — | 3.02x | 2.22x |
| table | narrow_100000 | projection | 1.25 MB | CPU | off | 885.6 μs | 847.5 μs | 2.94 ms | 4.31 ms | — | 3.46x | 5.09x |
| table | narrow_100000 | read_full | 1.25 MB | CPU | off | 1.22 ms | 1.13 ms | 3.14 ms | 1.11 ms | — | 2.79x | 0.99x |
| table | narrow_100000 | row_slice | 1.25 MB | CPU | off | 261.4 μs | 237.1 μs | 3.46 ms | 969.6 μs | — | 14.61x | 4.09x |
| table | narrow_100000 | scan_count | 1.25 MB | CPU | off | 42.7 μs | 44.7 μs | 745.9 μs | 99.6 μs | — | 17.49x | 2.34x |
| table | narrow_10000 | predicate_filter | 0.13 MB | CPU | off | 304.9 μs | 366.6 μs | 2.41 ms | 502.3 μs | — | 7.91x | 1.65x |
| table | narrow_10000 | predicate_filter_selective | 0.13 MB | CPU | off | 225.8 μs | 295.4 μs | 2.33 ms | 440.8 μs | — | 10.30x | 1.95x |
| table | narrow_10000 | projection | 0.13 MB | CPU | off | 228.5 μs | 187.8 μs | 2.32 ms | 651.0 μs | — | 12.35x | 3.47x |
| table | narrow_10000 | read_full | 0.13 MB | CPU | off | 258.9 μs | 225.2 μs | 2.36 ms | 306.1 μs | — | 10.47x | 1.36x |
| table | narrow_10000 | row_slice | 0.13 MB | CPU | off | 172.0 μs | 136.1 μs | 3.04 ms | 319.7 μs | — | 22.32x | 2.35x |
| table | narrow_10000 | scan_count | 0.13 MB | CPU | off | 42.2 μs | 41.0 μs | 733.9 μs | 104.8 μs | — | 17.89x | 2.55x |
| table | narrow_1000 | predicate_filter | 19.7 KB | CPU | off | 131.4 μs | 203.1 μs | 2.25 ms | 277.0 μs | — | 17.15x | 2.11x |
| table | narrow_1000 | predicate_filter_selective | 19.7 KB | CPU | off | 125.0 μs | 198.6 μs | 2.24 ms | 268.0 μs | — | 17.93x | 2.14x |
| table | narrow_1000 | projection | 19.7 KB | CPU | off | 170.7 μs | 182.6 μs | 3.43 ms | 424.8 μs | — | 20.12x | 2.49x |
| table | narrow_1000 | read_full | 19.7 KB | CPU | off | 175.1 μs | 145.8 μs | 2.30 ms | 239.5 μs | — | 15.79x | 1.64x |
| table | narrow_1000 | row_slice | 19.7 KB | CPU | off | 167.1 μs | 126.7 μs | 3.01 ms | 248.8 μs | — | 23.75x | 1.96x |
| table | narrow_1000 | scan_count | 19.7 KB | CPU | off | 42.4 μs | 46.2 μs | 762.0 μs | 111.8 μs | — | 17.97x | 2.64x |
| table | typed_100000 | predicate_filter | 2.39 MB | CPU | off | 841.1 μs | 851.5 μs | 1.86 ms | 1.36 ms | — | 2.21x | 1.61x |
| table | typed_100000 | predicate_filter_selective | 2.39 MB | CPU | off | 852.0 μs | 851.1 μs | 1.86 ms | 1.36 ms | — | 2.19x | 1.60x |
| table | typed_100000 | projection | 2.39 MB | CPU | off | 3.49 ms | 3.41 ms | 28.96 ms | 13.18 ms | — | 8.50x | 3.87x |
| table | typed_100000 | read_full | 2.39 MB | CPU | off | 5.11 ms | 5.10 ms | 29.04 ms | 14.09 ms | — | 5.69x | 2.76x |
| table | typed_100000 | row_slice | 2.39 MB | CPU | off | 620.7 μs | 596.9 μs | 4.72 ms | 1.91 ms | — | 7.92x | 3.20x |
| table | typed_100000 | scan_count | 2.39 MB | CPU | off | 30.8 μs | 29.0 μs | 445.7 μs | 67.9 μs | — | 15.39x | 2.35x |
| table | typed_10000 | predicate_filter | 0.24 MB | CPU | off | 178.9 μs | 213.2 μs | 1.43 ms | 279.7 μs | — | 7.99x | 1.56x |
| table | typed_10000 | predicate_filter_selective | 0.24 MB | CPU | off | 176.3 μs | 222.1 μs | 1.44 ms | 278.9 μs | — | 8.17x | 1.58x |
| table | typed_10000 | projection | 0.24 MB | CPU | off | 454.2 μs | 423.7 μs | 4.08 ms | 1.46 ms | — | 9.63x | 3.45x |
| table | typed_10000 | read_full | 0.24 MB | CPU | off | 616.5 μs | 597.2 μs | 4.12 ms | 1.59 ms | — | 6.90x | 2.67x |
| table | typed_10000 | row_slice | 0.24 MB | CPU | off | 160.4 μs | 137.2 μs | 2.17 ms | 389.2 μs | — | 15.85x | 2.84x |
| table | typed_10000 | scan_count | 0.24 MB | CPU | off | 28.1 μs | 26.4 μs | 415.8 μs | 64.5 μs | — | 15.77x | 2.45x |
| table | varlen_100000 | predicate_filter | 3.06 MB | CPU | off | 870.5 μs | 732.3 μs | 1.73 ms | 1.23 ms | — | 2.37x | 1.68x |
| table | varlen_100000 | predicate_filter_selective | 3.06 MB | CPU | off | 695.0 μs | 724.4 μs | 1.71 ms | 1.24 ms | — | 2.45x | 1.78x |
| table | varlen_100000 | projection | 3.06 MB | CPU | off | 72.73 ms | 10.41 ms | 527.03 ms | 106.65 ms | — | 50.64x | 10.25x |
| table | varlen_100000 | read_full | 3.06 MB | CPU | off | 74.74 ms | 10.80 ms | 528.67 ms | 107.37 ms | — | 48.93x | 9.94x |
| table | varlen_100000 | row_slice | 3.06 MB | CPU | off | 6.96 ms | 1.04 ms | 54.18 ms | 11.73 ms | — | 51.96x | 11.25x |
| table | varlen_100000 | scan_count | 3.06 MB | CPU | off | 28.1 μs | 27.9 μs | 447.1 μs | 64.0 μs | — | 16.03x | 2.30x |
| table | varlen_10000 | predicate_filter | 0.31 MB | CPU | off | 165.9 μs | 216.0 μs | 1.33 ms | 269.7 μs | — | 8.00x | 1.62x |
| table | varlen_10000 | predicate_filter_selective | 0.31 MB | CPU | off | 155.1 μs | 211.4 μs | 1.34 ms | 265.6 μs | — | 8.62x | 1.71x |
| table | varlen_10000 | projection | 0.31 MB | CPU | off | 6.97 ms | 1.05 ms | 53.21 ms | 10.73 ms | — | 50.56x | 10.19x |
| table | varlen_10000 | read_full | 0.31 MB | CPU | off | 6.98 ms | 1.05 ms | 53.16 ms | 10.76 ms | — | 50.49x | 10.22x |
| table | varlen_10000 | row_slice | 0.31 MB | CPU | off | 808.0 μs | 184.5 μs | 6.98 ms | 1.36 ms | — | 37.82x | 7.35x |
| table | varlen_10000 | scan_count | 0.31 MB | CPU | off | 28.5 μs | 26.5 μs | 428.5 μs | 60.1 μs | — | 16.16x | 2.27x |
| table | varlen_1000 | predicate_filter | 39.4 KB | CPU | off | 70.8 μs | 129.3 μs | 1.28 ms | 161.7 μs | — | 18.09x | 2.28x |
| table | varlen_1000 | predicate_filter_selective | 39.4 KB | CPU | off | 72.3 μs | 129.6 μs | 1.28 ms | 166.5 μs | — | 17.65x | 2.30x |
| table | varlen_1000 | projection | 39.4 KB | CPU | off | 793.4 μs | 195.3 μs | 6.58 ms | 1.25 ms | — | 33.72x | 6.43x |
| table | varlen_1000 | read_full | 39.4 KB | CPU | off | 801.8 μs | 193.1 μs | 6.59 ms | 1.24 ms | — | 34.10x | 6.44x |
| table | varlen_1000 | row_slice | 39.4 KB | CPU | off | 197.9 μs | 102.2 μs | 2.22 ms | 296.3 μs | — | 21.76x | 2.90x |
| table | varlen_1000 | scan_count | 39.4 KB | CPU | off | 25.9 μs | 25.3 μs | 425.0 μs | 61.1 μs | — | 16.81x | 2.42x |
| table | wide_100000 | predicate_filter | 20.71 MB | CPU | off | 3.43 ms | 3.20 ms | 8.31 ms | 4.17 ms | — | 2.60x | 1.31x |
| table | wide_100000 | predicate_filter_selective | 20.71 MB | CPU | off | 3.03 ms | 3.03 ms | 7.91 ms | 3.84 ms | — | 2.61x | 1.27x |
| table | wide_100000 | projection | 20.71 MB | CPU | off | 4.70 ms | 4.63 ms | 14.14 ms | 8.61 ms | — | 3.05x | 1.86x |
| table | wide_100000 | read_full | 20.71 MB | CPU | off | 31.16 ms | 31.17 ms | 227.80 ms | 70.80 ms | — | 7.31x | 2.27x |
| table | wide_100000 | row_slice | 20.71 MB | CPU | off | 2.11 ms | 1.41 ms | 21.57 ms | 4.98 ms | — | 15.28x | 3.53x |
| table | wide_100000 | scan_count | 20.71 MB | CPU | off | 38.6 μs | 37.9 μs | 559.9 μs | 256.0 μs | — | 14.76x | 6.75x |
| table | wide_10000 | predicate_filter | 2.08 MB | CPU | off | 789.2 μs | 868.4 μs | 10.64 ms | 1.27 ms | — | 13.48x | 1.60x |
| table | wide_10000 | predicate_filter_selective | 2.08 MB | CPU | off | 693.3 μs | 746.9 μs | 10.12 ms | 1.20 ms | — | 14.60x | 1.73x |
| table | wide_10000 | projection | 2.08 MB | CPU | off | 429.4 μs | 403.5 μs | 5.93 ms | 871.7 μs | — | 14.71x | 2.16x |
| table | wide_10000 | read_full | 2.08 MB | CPU | off | 1.48 ms | 1.48 ms | 17.15 ms | 4.52 ms | — | 11.59x | 3.05x |
| table | wide_10000 | row_slice | 2.08 MB | CPU | off | 850.6 μs | 856.6 μs | 18.23 ms | 1.50 ms | — | 21.43x | 1.76x |
| table | wide_10000 | scan_count | 2.08 MB | CPU | off | 65.1 μs | 63.3 μs | 943.9 μs | 426.7 μs | — | 14.91x | 6.74x |
| table | wide_1000 | predicate_filter | 0.22 MB | CPU | off | 233.3 μs | 288.1 μs | 9.72 ms | 651.7 μs | — | 41.64x | 2.79x |
| table | wide_1000 | predicate_filter_selective | 0.22 MB | CPU | off | 213.2 μs | 278.8 μs | 9.75 ms | 642.1 μs | — | 45.75x | 3.01x |
| table | wide_1000 | projection | 0.22 MB | CPU | off | 260.8 μs | 225.7 μs | 9.76 ms | 656.1 μs | — | 43.25x | 2.91x |
| table | wide_1000 | read_full | 0.22 MB | CPU | off | 830.7 μs | 854.6 μs | 12.30 ms | 1.38 ms | — | 14.81x | 1.67x |
| table | wide_1000 | row_slice | 0.22 MB | CPU | off | 732.1 μs | 741.6 μs | 16.24 ms | 842.4 μs | — | 22.19x | 1.15x |
| table | wide_1000 | scan_count | 0.22 MB | CPU | off | 60.4 μs | 62.5 μs | 942.1 μs | 423.5 μs | — | 15.59x | 7.01x |
| table | ascii_10000 | predicate_filter | 0.44 MB | CPU | on | 411.7 μs | 417.2 μs | 2.66 ms | — | — | 6.46x | — |
| table | ascii_10000 | predicate_filter_selective | 0.44 MB | CPU | on | 402.6 μs | 409.3 μs | 2.67 ms | — | — | 6.63x | — |
| table | ascii_10000 | projection | 0.44 MB | CPU | on | 1.08 ms | 1.07 ms | 8.03 ms | — | — | 7.49x | — |
| table | ascii_10000 | read_full | 0.44 MB | CPU | on | 1.08 ms | 1.06 ms | 8.01 ms | — | — | 7.56x | — |
| table | ascii_10000 | row_slice | 0.44 MB | CPU | on | 236.1 μs | 217.8 μs | 2.57 ms | — | — | 11.78x | — |
| table | ascii_10000 | scan_count | 0.44 MB | CPU | on | 24.5 μs | 24.8 μs | 407.3 μs | — | — | 16.64x | — |
| table | ascii_1000 | predicate_filter | 50.6 KB | CPU | on | 177.7 μs | 189.6 μs | 1.51 ms | — | — | 8.52x | — |
| table | ascii_1000 | predicate_filter_selective | 50.6 KB | CPU | on | 184.8 μs | 181.6 μs | 1.51 ms | — | — | 8.33x | — |
| table | ascii_1000 | projection | 50.6 KB | CPU | on | 227.2 μs | 217.4 μs | 2.18 ms | — | — | 10.03x | — |
| table | ascii_1000 | read_full | 50.6 KB | CPU | on | 233.5 μs | 212.1 μs | 2.19 ms | — | — | 10.34x | — |
| table | ascii_1000 | row_slice | 50.6 KB | CPU | on | 162.7 μs | 135.2 μs | 1.94 ms | — | — | 14.38x | — |
| table | ascii_1000 | scan_count | 50.6 KB | CPU | on | 27.7 μs | 30.7 μs | 486.7 μs | — | — | 17.58x | — |
| table | mixed_1000000 | predicate_filter | 50.55 MB | CPU | on | 5.40 ms | 4.95 ms | 11.96 ms | — | — | 2.42x | — |
| table | mixed_1000000 | predicate_filter_selective | 50.55 MB | CPU | on | 4.15 ms | 4.18 ms | 8.67 ms | — | — | 2.09x | — |
| table | mixed_1000000 | projection | 50.55 MB | CPU | on | 7.33 ms | 7.95 ms | 12.92 ms | — | — | 1.76x | — |
| table | mixed_1000000 | read_full | 50.55 MB | CPU | on | 28.77 ms | 25.37 ms | 337.72 ms | — | — | 13.31x | — |
| table | mixed_1000000 | row_slice | 50.55 MB | CPU | on | 276.5 μs | 217.8 μs | 9.07 ms | — | — | 41.66x | — |
| table | mixed_1000000 | scan_count | 50.55 MB | CPU | on | 33.4 μs | 31.2 μs | 466.8 μs | — | — | 14.98x | — |
| table | mixed_100000 | predicate_filter | 5.06 MB | CPU | on | 1.25 ms | 1.02 ms | 4.83 ms | — | — | 4.74x | — |
| table | mixed_100000 | predicate_filter_selective | 5.06 MB | CPU | on | 878.9 μs | 873.7 μs | 4.27 ms | — | — | 4.89x | — |
| table | mixed_100000 | projection | 5.06 MB | CPU | on | 1.31 ms | 1.21 ms | 4.83 ms | — | — | 3.99x | — |
| table | mixed_100000 | read_full | 5.06 MB | CPU | on | 2.77 ms | 2.68 ms | 52.13 ms | — | — | 19.49x | — |
| table | mixed_100000 | row_slice | 5.06 MB | CPU | on | 422.5 μs | 328.6 μs | 9.68 ms | — | — | 29.47x | — |
| table | mixed_100000 | scan_count | 5.06 MB | CPU | on | 48.8 μs | 49.2 μs | 794.2 μs | — | — | 16.29x | — |
| table | mixed_10000 | predicate_filter | 0.51 MB | CPU | on | 313.8 μs | 349.2 μs | 3.30 ms | — | — | 10.51x | — |
| table | mixed_10000 | predicate_filter_selective | 0.51 MB | CPU | on | 225.6 μs | 267.5 μs | 3.24 ms | — | — | 14.37x | — |
| table | mixed_10000 | projection | 0.51 MB | CPU | on | 269.3 μs | 192.0 μs | 3.26 ms | — | — | 16.99x | — |
| table | mixed_10000 | read_full | 0.51 MB | CPU | on | 410.3 μs | 316.9 μs | 7.83 ms | — | — | 24.70x | — |
| table | mixed_10000 | row_slice | 0.51 MB | CPU | on | 260.1 μs | 180.9 μs | 5.07 ms | — | — | 28.05x | — |
| table | mixed_10000 | scan_count | 0.51 MB | CPU | on | 47.9 μs | 46.5 μs | 759.6 μs | — | — | 16.33x | — |
| table | mixed_1000 | predicate_filter | 0.06 MB | CPU | on | 142.9 μs | 195.3 μs | 3.17 ms | — | — | 22.17x | — |
| table | mixed_1000 | predicate_filter_selective | 0.06 MB | CPU | on | 283.2 μs | 193.4 μs | 3.13 ms | — | — | 16.17x | — |
| table | mixed_1000 | projection | 0.06 MB | CPU | on | 216.5 μs | 142.5 μs | 3.13 ms | — | — | 21.97x | — |
| table | mixed_1000 | read_full | 0.06 MB | CPU | on | 258.0 μs | 179.2 μs | 3.76 ms | — | — | 20.99x | — |
| table | mixed_1000 | row_slice | 0.06 MB | CPU | on | 251.5 μs | 162.1 μs | 4.63 ms | — | — | 28.55x | — |
| table | mixed_1000 | scan_count | 0.06 MB | CPU | on | 45.8 μs | 48.1 μs | 785.4 μs | — | — | 17.14x | — |
| table | narrow_1000000 | predicate_filter | 12.40 MB | CPU | on | 3.07 ms | 2.40 ms | 8.00 ms | — | — | 3.34x | — |
| table | narrow_1000000 | predicate_filter_selective | 12.40 MB | CPU | on | 1.61 ms | 1.60 ms | 4.66 ms | — | — | 2.91x | — |
| table | narrow_1000000 | projection | 12.40 MB | CPU | on | 2.01 ms | 1.97 ms | 5.05 ms | — | — | 2.56x | — |
| table | narrow_1000000 | read_full | 12.40 MB | CPU | on | 3.13 ms | 3.05 ms | 6.28 ms | — | — | 2.06x | — |
| table | narrow_1000000 | row_slice | 12.40 MB | CPU | on | 172.8 μs | 118.4 μs | 3.42 ms | — | — | 28.90x | — |
| table | narrow_1000000 | scan_count | 12.40 MB | CPU | on | 25.6 μs | 27.3 μs | 437.4 μs | — | — | 17.07x | — |
| table | narrow_100000 | predicate_filter | 1.25 MB | CPU | on | 799.7 μs | 689.2 μs | 3.49 ms | — | — | 5.06x | — |
| table | narrow_100000 | predicate_filter_selective | 1.25 MB | CPU | on | 460.0 μs | 464.2 μs | 2.90 ms | — | — | 6.29x | — |
| table | narrow_100000 | projection | 1.25 MB | CPU | on | 478.1 μs | 380.9 μs | 2.89 ms | — | — | 7.58x | — |
| table | narrow_100000 | read_full | 1.25 MB | CPU | on | 661.5 μs | 585.9 μs | 3.10 ms | — | — | 5.28x | — |
| table | narrow_100000 | row_slice | 1.25 MB | CPU | on | 255.7 μs | 170.8 μs | 3.41 ms | — | — | 19.99x | — |
| table | narrow_100000 | scan_count | 1.25 MB | CPU | on | 44.1 μs | 49.1 μs | 757.8 μs | — | — | 17.18x | — |
| table | narrow_10000 | predicate_filter | 0.13 MB | CPU | on | 269.4 μs | 320.4 μs | 2.42 ms | — | — | 8.97x | — |
| table | narrow_10000 | predicate_filter_selective | 0.13 MB | CPU | on | 184.7 μs | 235.5 μs | 2.34 ms | — | — | 12.65x | — |
| table | narrow_10000 | projection | 0.13 MB | CPU | on | 223.6 μs | 137.8 μs | 2.34 ms | — | — | 16.96x | — |
| table | narrow_10000 | read_full | 0.13 MB | CPU | on | 250.7 μs | 165.3 μs | 2.40 ms | — | — | 14.55x | — |
| table | narrow_10000 | row_slice | 0.13 MB | CPU | on | 210.5 μs | 128.1 μs | 3.05 ms | — | — | 23.82x | — |
| table | narrow_10000 | scan_count | 0.13 MB | CPU | on | 45.9 μs | 43.3 μs | 757.7 μs | — | — | 17.51x | — |
| table | narrow_1000 | predicate_filter | 19.7 KB | CPU | on | 124.0 μs | 198.7 μs | 2.27 ms | — | — | 18.29x | — |
| table | narrow_1000 | predicate_filter_selective | 19.7 KB | CPU | on | 127.9 μs | 192.2 μs | 2.28 ms | — | — | 17.84x | — |
| table | narrow_1000 | projection | 19.7 KB | CPU | on | 198.3 μs | 124.9 μs | 2.28 ms | — | — | 18.24x | — |
| table | narrow_1000 | read_full | 19.7 KB | CPU | on | 201.4 μs | 124.0 μs | 2.32 ms | — | — | 18.72x | — |
| table | narrow_1000 | row_slice | 19.7 KB | CPU | on | 205.5 μs | 131.7 μs | 3.03 ms | — | — | 23.04x | — |
| table | narrow_1000 | scan_count | 19.7 KB | CPU | on | 45.3 μs | 47.7 μs | 776.3 μs | — | — | 17.14x | — |
| table | typed_100000 | predicate_filter | 2.39 MB | CPU | on | 614.7 μs | 476.2 μs | 1.75 ms | — | — | 3.68x | — |
| table | typed_100000 | predicate_filter_selective | 2.39 MB | CPU | on | 442.8 μs | 465.3 μs | 1.79 ms | — | — | 4.05x | — |
| table | typed_100000 | projection | 2.39 MB | CPU | on | 1.23 ms | 1.17 ms | 28.68 ms | — | — | 24.62x | — |
| table | typed_100000 | read_full | 2.39 MB | CPU | on | 1.36 ms | 1.26 ms | 28.93 ms | — | — | 22.95x | — |
| table | typed_100000 | row_slice | 2.39 MB | CPU | on | 266.0 μs | 210.3 μs | 4.48 ms | — | — | 21.31x | — |
| table | typed_100000 | scan_count | 2.39 MB | CPU | on | 30.3 μs | 27.3 μs | 442.3 μs | — | — | 16.18x | — |
| table | typed_10000 | predicate_filter | 0.24 MB | CPU | on | 154.9 μs | 189.9 μs | 1.44 ms | — | — | 9.26x | — |
| table | typed_10000 | predicate_filter_selective | 0.24 MB | CPU | on | 152.9 μs | 181.8 μs | 1.43 ms | — | — | 9.36x | — |
| table | typed_10000 | projection | 0.24 MB | CPU | on | 249.2 μs | 198.7 μs | 4.03 ms | — | — | 20.29x | — |
| table | typed_10000 | read_full | 0.24 MB | CPU | on | 257.7 μs | 202.7 μs | 4.08 ms | — | — | 20.14x | — |
| table | typed_10000 | row_slice | 0.24 MB | CPU | on | 149.0 μs | 95.5 μs | 2.17 ms | — | — | 22.73x | — |
| table | typed_10000 | scan_count | 0.24 MB | CPU | on | 27.9 μs | 27.8 μs | 432.7 μs | — | — | 15.55x | — |
| table | varlen_100000 | predicate_filter | 3.06 MB | CPU | on | 578.4 μs | 428.0 μs | 1.58 ms | — | — | 3.69x | — |
| table | varlen_100000 | predicate_filter_selective | 3.06 MB | CPU | on | 389.5 μs | 398.7 μs | 1.60 ms | — | — | 4.10x | — |
| table | varlen_100000 | projection | 3.06 MB | CPU | on | 74.18 ms | 10.94 ms | 530.40 ms | — | — | 48.46x | — |
| table | varlen_100000 | read_full | 3.06 MB | CPU | on | 75.00 ms | 10.93 ms | 531.59 ms | — | — | 48.64x | — |
| table | varlen_100000 | row_slice | 3.06 MB | CPU | on | 7.07 ms | 1.12 ms | 54.19 ms | — | — | 48.31x | — |
| table | varlen_100000 | scan_count | 3.06 MB | CPU | on | 28.1 μs | 28.2 μs | 446.6 μs | — | — | 15.91x | — |
| table | varlen_10000 | predicate_filter | 0.31 MB | CPU | on | 152.2 μs | 183.4 μs | 1.32 ms | — | — | 8.67x | — |
| table | varlen_10000 | predicate_filter_selective | 0.31 MB | CPU | on | 149.9 μs | 175.7 μs | 1.31 ms | — | — | 8.76x | — |
| table | varlen_10000 | projection | 0.31 MB | CPU | on | 7.07 ms | 1.14 ms | 53.80 ms | — | — | 47.28x | — |
| table | varlen_10000 | read_full | 0.31 MB | CPU | on | 7.05 ms | 1.15 ms | 54.03 ms | — | — | 47.02x | — |
| table | varlen_10000 | row_slice | 0.31 MB | CPU | on | 866.5 μs | 284.1 μs | 7.02 ms | — | — | 24.69x | — |
| table | varlen_10000 | scan_count | 0.31 MB | CPU | on | 25.9 μs | 26.7 μs | 428.1 μs | — | — | 16.56x | — |
| table | varlen_1000 | predicate_filter | 39.4 KB | CPU | on | 82.1 μs | 130.6 μs | 1.27 ms | — | — | 15.52x | — |
| table | varlen_1000 | predicate_filter_selective | 39.4 KB | CPU | on | 79.3 μs | 126.8 μs | 1.28 ms | — | — | 16.17x | — |
| table | varlen_1000 | projection | 39.4 KB | CPU | on | 861.0 μs | 280.6 μs | 6.62 ms | — | — | 23.59x | — |
| table | varlen_1000 | read_full | 39.4 KB | CPU | on | 870.5 μs | 238.9 μs | 6.60 ms | — | — | 27.64x | — |
| table | varlen_1000 | row_slice | 39.4 KB | CPU | on | 262.0 μs | 148.8 μs | 2.24 ms | — | — | 15.04x | — |
| table | varlen_1000 | scan_count | 39.4 KB | CPU | on | 26.6 μs | 26.8 μs | 424.9 μs | — | — | 15.96x | — |
| table | wide_100000 | predicate_filter | 20.71 MB | CPU | on | 1.58 ms | 1.37 ms | 7.36 ms | — | — | 5.38x | — |
| table | wide_100000 | predicate_filter_selective | 20.71 MB | CPU | on | 1.22 ms | 1.24 ms | 7.06 ms | — | — | 5.81x | — |
| table | wide_100000 | projection | 20.71 MB | CPU | on | 1.67 ms | 1.60 ms | 7.58 ms | — | — | 4.73x | — |
| table | wide_100000 | read_full | 20.71 MB | CPU | on | 24.64 ms | 13.77 ms | 128.56 ms | — | — | 9.34x | — |
| table | wide_100000 | row_slice | 20.71 MB | CPU | on | 1.03 ms | 977.6 μs | 20.76 ms | — | — | 21.24x | — |
| table | wide_100000 | scan_count | 20.71 MB | CPU | on | 36.9 μs | 36.9 μs | 565.5 μs | — | — | 15.32x | — |
| table | wide_10000 | predicate_filter | 2.08 MB | CPU | on | 483.9 μs | 510.7 μs | 10.02 ms | — | — | 20.70x | — |
| table | wide_10000 | predicate_filter_selective | 2.08 MB | CPU | on | 398.8 μs | 429.5 μs | 9.95 ms | — | — | 24.94x | — |
| table | wide_10000 | projection | 2.08 MB | CPU | on | 451.2 μs | 370.8 μs | 9.98 ms | — | — | 26.91x | — |
| table | wide_10000 | read_full | 2.08 MB | CPU | on | 1.64 ms | 1.53 ms | 28.54 ms | — | — | 18.67x | — |
| table | wide_10000 | row_slice | 2.08 MB | CPU | on | 837.8 μs | 759.7 μs | 17.93 ms | — | — | 23.60x | — |
| table | wide_10000 | scan_count | 2.08 MB | CPU | on | 62.7 μs | 62.1 μs | 973.5 μs | — | — | 15.67x | — |
| table | wide_1000 | predicate_filter | 0.22 MB | CPU | on | 197.0 μs | 252.3 μs | 9.65 ms | — | — | 48.97x | — |
| table | wide_1000 | predicate_filter_selective | 0.22 MB | CPU | on | 193.8 μs | 251.3 μs | 9.69 ms | — | — | 50.00x | — |
| table | wide_1000 | projection | 0.22 MB | CPU | on | 280.8 μs | 207.2 μs | 9.64 ms | — | — | 46.55x | — |
| table | wide_1000 | read_full | 0.22 MB | CPU | on | 832.6 μs | 745.1 μs | 12.23 ms | — | — | 16.41x | — |
| table | wide_1000 | row_slice | 0.22 MB | CPU | on | 760.6 μs | 695.5 μs | 16.13 ms | — | — | 23.19x | — |
| table | wide_1000 | scan_count | 0.22 MB | CPU | on | 62.6 μs | 62.6 μs | 985.6 μs | — | — | 15.75x | — |
Performance comparisons & edge cases¶
Cases where torchfits is not first in its comparison family (CPU and GPU). GPU lags may reflect software or hardware limits — they are listed, not hidden.
| Platform | Domain | Case | mmap | torchfits | Peak RSS (MB) | Winner | Lag |
|---|---|---|---|---|---|---|---|
| Linux x86_64 / CPU | tensor | compressed_hcompress_1 [read_full] | on | 45.70 ms | 293.8 | fitsio/fitsio_torch | 1.02× |
| Linux x86_64 / CPU | tensor | compressed_hcompress_1 [read_full] | off | 45.56 ms | 309.3 | fitsio/fitsio_torch | 1.02× |
| Linux x86_64 / CPU | tensor | compressed_hcompress_1 [read_full] | on | 45.72 ms | 293.8 | fitsio/fitsio_torch | 1.02× |
| Linux x86_64 / CPU | tensor | compressed_hcompress_1 [read_full] | off | 45.54 ms | 309.3 | fitsio/fitsio_torch | 1.02× |
| Linux x86_64 / CPU | table | narrow_100000 [read_full] | off | 1.22 ms | 380.4 | fitsio/fitsio_torch | 1.36× |
| Linux x86_64 / CPU | table | narrow_1000000 [read_full] | off | 6.00 ms | 399.4 | fitsio/fitsio_torch | 1.21× |
| Linux x86_64 / CPU | table | ascii_10000 [predicate_filter] | off | 382.6 μs | 422.3 | fitsio/fitsio_torch | 1.02× |
| Linux x86_64 / CPU | table | ascii_10000 [predicate_filter] | off | 379.7 μs | 422.3 | fitsio/fitsio | 1.06× |
| Linux x86_64 / CPU | table | ascii_10000 [predicate_filter_selective] | off | 371.9 μs | 422.3 | fitsio/fitsio | 1.03× |
| Linux x86_64 / CPU | table | narrow_100000 [read_full] | off | 1.13 ms | 380.4 | fitsio/fitsio | 1.01× |
| Linux x86_64 / CPU | table | narrow_1000000 [read_full] | off | 5.93 ms | 403.5 | fitsio/fitsio | 1.01× |
| Linux x86_64 / CUDA | tensor | tiny_int8_1d [read_full @ cuda] | off | 118.1 μs | 766.8 | fitsio/fitsio_torch_device | 1.10× |
| Linux x86_64 / CUDA | tensor | tiny_float64_3d [read_full @ cuda] | off | 131.0 μs | 766.8 | fitsio/fitsio_torch_device | 1.08× |
| Linux x86_64 / CUDA | tensor | tiny_int16_2d [read_full @ cuda] | off | 117.4 μs | 766.8 | fitsio/fitsio_torch_device | 1.06× |
| Linux x86_64 / CUDA | tensor | tiny_int64_1d [read_full @ cuda] | off | 111.0 μs | 766.8 | fitsio/fitsio_torch_device | 1.05× |
| Linux x86_64 / CUDA | tensor | compressed_hcompress_1 [read_full @ cuda] | off | 30.72 ms | 729.4 | fitsio/fitsio_torch_device | 1.04× |
| Linux x86_64 / CUDA | tensor | medium_int8_1d [read_full @ cuda] | off | 145.8 μs | 766.8 | fitsio/fitsio_torch_device | 1.04× |
| Linux x86_64 / CUDA | tensor | compressed_hcompress_1 [read_full @ cuda] | on | 30.60 ms | 699.1 | fitsio/fitsio_torch_device | 1.03× |
| Linux x86_64 / CUDA | tensor | tiny_int64_2d [read_full @ cuda] | off | 125.7 μs | 766.8 | fitsio/fitsio_torch_device | 1.03× |
| Linux x86_64 / CUDA | tensor | tiny_int8_2d [read_full @ cuda] | off | 117.5 μs | 766.8 | fitsio/fitsio_torch_device | 1.03× |
| Linux x86_64 / CUDA | tensor | compressed_hcompress_1 [read_full] | on | 30.27 ms | 606.7 | fitsio/fitsio_torch | 1.03× |
| Linux x86_64 / CUDA | tensor | tiny_float64_1d [read_full @ cuda] | off | 112.6 μs | 766.8 | fitsio/fitsio_torch_device | 1.02× |
| Linux x86_64 / CUDA | tensor | small_int8_1d [read_full @ cuda] | off | 113.2 μs | 766.8 | fitsio/fitsio_torch_device | 1.02× |
| Linux x86_64 / CUDA | tensor | compressed_hcompress_1 [read_full] | off | 30.24 ms | 728.3 | fitsio/fitsio_torch | 1.02× |
| Linux x86_64 / CUDA | tensor | small_uint16_2d [read_full @ cuda] | off | 153.5 μs | 766.8 | fitsio/fitsio_torch_device | 1.02× |
| Linux x86_64 / CUDA | tensor | tiny_float64_2d [read_full @ cuda] | off | 119.2 μs | 766.8 | fitsio/fitsio_torch_device | 1.01× |
| Linux x86_64 / CUDA | tensor | small_int64_1d [read_full @ cuda] | off | 133.1 μs | 766.8 | fitsio/fitsio_torch_device | 1.01× |
| Linux x86_64 / CUDA | tensor | medium_int8_2d [read_full] | off | 321.0 μs | 765.8 | fitsio/fitsio_torch | 1.01× |
| Linux x86_64 / CUDA | tensor | tiny_int32_2d [read_full @ cuda] | off | 114.4 μs | 766.8 | fitsio/fitsio_torch_device | 1.01× |
| Linux x86_64 / CUDA | tensor | compressed_hcompress_1 [read_full @ cuda] | off | 30.73 ms | 729.4 | fitsio/fitsio_torch_device_specialized | 1.04× |
| Linux x86_64 / CUDA | tensor | compressed_hcompress_1 [read_full @ cuda] | on | 30.58 ms | 699.1 | fitsio/fitsio_torch_device_specialized | 1.03× |
| Linux x86_64 / CUDA | tensor | compressed_hcompress_1 [read_full] | on | 30.31 ms | 606.7 | fitsio/fitsio_torch | 1.03× |
| Linux x86_64 / CUDA | tensor | compressed_hcompress_1 [read_full] | off | 30.33 ms | 728.3 | fitsio/fitsio_torch | 1.02× |
| Linux x86_64 / CUDA | tensor | small_int16_3d [read_full] | off | 156.1 μs | 765.8 | fitsio/fitsio_torch | 1.02× |
| Linux x86_64 / CUDA | table | narrow_1000000 [read_full] | off | 7.68 ms | 715.5 | fitsio/fitsio_torch | 1.08× |
| Linux x86_64 / CUDA | table | typed_100000 [predicate_filter] | off | 2.09 ms | 738.1 | fitsio/fitsio_torch | 1.03× |
| Linux x86_64 / CUDA | table | typed_100000 [predicate_filter] | off | 2.10 ms | 738.1 | fitsio/fitsio | 1.03× |
Published runs by platform¶
| Platform | Run ID | Rows | Time deficits | Median peak RSS (MB) | Notes |
|---|---|---|---|---|---|
| Linux x86_64 / CPU | exhaustive_cpu_20260807_013736 | 3057 | 11 | 293.8 | lab + mmap-matrix |
| Linux x86_64 / CUDA | exhaustive_cuda_20260807_013736 | 4315 | 26 | 719.3 | lab + mmap-matrix + GPU |
Historical July 2026 runs: MPS exhaustive_mps_20260719_143706 (local);
CANFAR staging CPU exhaustive_cpu_20260719_144337 and CUDA
exhaustive_cuda_20260719_144457 (clone bench/thin-io-scorecard @ 9b9e7cf).
ML loader: ml_20260719_145743. MegaCam: 20260719_075555.
Latest local quick benchmark evidence:
| Run ID | Scope | Command | Rows | Deficits |
|---|---|---|---|---|
| — | FITS image I/O | (no run yet) | — | — |
| — | FITS table I/O | (no run yet) | — | — |
ML DataLoader throughput¶
Run pixi run bench-ml to populate ML loader throughput.
CFHT MegaCam MEF cutouts (local)¶
Source: docs/assets/bench/20260719_075555/megacam_results.csv (160 OK rows).
Median throughput over OK rows (earlier table values were copy-paste μs from unrelated suites).
| Method | Median throughput |
|---|---|
fitsio_cached |
52.7 MB/s |
torchfits_cached |
49.3 MB/s |
torchfits_materialize |
119.4 MB/s |
torchfits_naive |
50.6 MB/s |
Keep this page current with the latest tensor and table benchmark run before making performance claims.