API surface, CLI commands, and operator tables. Look here for "what exists and what it accepts." For task-oriented usage see how-to.md; for design rationale see explanation.md.
- Core types
- Reader API
- Writer API
- Scan API
- Encoding registry
- FSST (
io.github.dfa1.vortex.fsst) - Parquet / CSV import and Parquet export
- JDBC import
- Calcite SQL adapter
- CLI
- Encoding compatibility
Physical primitive type — wire-level numeric kind for a column.
| Constant | Bytes | Notes |
|---|---|---|
U8, U16, U32, U64 |
1 / 2 / 4 / 8 | Unsigned integers |
I8, I16, I32, I64 |
1 / 2 / 4 / 8 | Signed integers |
F16 |
2 | IEEE 754 half — decoded via Float#float16ToFloat |
F32, F64 |
4 / 8 | IEEE 754 single / double |
Methods: byteSize(), isFloating(), isInteger(), isSigned().
Sealed logical type. All variants take a trailing boolean nullable.
Each DType variant is a record. Prefer the canonical constants / factories for
new code — the record constructors stay available for pattern matching and tests.
| Record | Constant / factory | Record constructor |
|---|---|---|
DType.Null |
DType.NULL |
new DType.Null(nullable) |
DType.Bool |
DType.BOOL |
new DType.Bool(nullable) |
DType.Primitive |
DType.I8 … DType.I64, DType.U8 … DType.U64, DType.F16, DType.F32, DType.F64 (constants) |
DType.I64.asNullable() / new DType.Primitive(PType, nullable) |
DType.Decimal |
DType.decimal(precision, scale) |
new DType.Decimal(precision, scale, nullable) |
DType.Utf8 |
DType.UTF8 |
new DType.Utf8(nullable) |
DType.Binary |
DType.BINARY |
new DType.Binary(nullable) |
DType.Variant |
DType.VARIANT |
new DType.Variant(nullable) |
DType.Struct |
DType.structBuilder().field(name, type)…build() |
new DType.Struct(fieldNames, fieldTypes, nullable) |
DType.List |
— | new DType.List(elementType, nullable) |
DType.FixedSizeList |
— | new DType.FixedSizeList(elementType, fixedSize, nullable) |
DType.Map |
— | new DType.Map(keyType, valueType, keysSorted, nullable) |
DType.Extension |
— | new DType.Extension(id, storageDType, metadata, nullable) |
DType.Union |
— (read only: no union is written) | new DType.Union(names, variantTypes, typeIds, nullable); variantIndex(typeId) |
Helpers: nullable() (boolean accessor on every record), asNullable() (fluent
shortcut returning a nullable copy), withNullable(boolean), DType.Struct.field(name),
DType.Map.entriesDtype() (the non-nullable {key, value} struct backing a map). A map's
keyType must be non-nullable — new DType.Map(...) throws VortexException otherwise.
| Type | Shape | Notes |
|---|---|---|
EncodingId |
sealed interface — WellKnown enum + Custom record |
Array-encoding identity; total parse(String) over non-blank ids; constants re-exported (EncodingId.VORTEX_PRIMITIVE, …) |
LayoutId |
sealed interface — WellKnown enum + Custom record |
Layout identity (separate namespace from encodings; vortex.flat is layout-only); both zoned aliases vortex.zoned/vortex.stats |
ColumnName |
record ColumnName(String value) |
Validated column name: non-blank, no control characters. ColumnName.violation(String) is the policy chokepoint shared by builder, writer, and file parser |
MemorySize |
record MemorySize(long bytes) |
Validated non-negative byte count; ofKiB/ofMiB/ofGiB factories, toGiB() for display. Rejects negative values (IllegalArgumentException) — a programmer-error guard, not a VortexException-worthy untrusted-input check |
EditionFamily |
enum — CORE, PREVIEW, ZSTD |
Closed (unlike EncodingId/LayoutId): an edition family is a cross-implementation compatibility promise, so a custom family carries no real guarantee — see Editions |
EditionId |
record EditionId(EditionFamily family, YearMonth cutMonth, int version) |
e.g. core2025.05.0; isAtOrBefore orders editions within a family only |
Edition |
record Edition(EditionId id, Set<EncodingId> added) |
Only Editions's catalog constants should be constructed (API contract, not compiler-enforced — see ADR 0023) |
Interface common to both handle implementations below — everything scan-shaped
(dtype(), layout(), scan(ScanOptions), rawSegment(SegmentSpec), …) is declared here so code
that only reads (e.g. ParquetExporter.exportParquet(VortexHandle, Path)) doesn't care whether the
source is a local file or a remote URL. Implements Closeable.
| Implementation | Backing storage |
|---|---|
VortexReader |
Memory-mapped local file |
VortexHttpReader |
Remote file read via HTTP Range requests |
Memory-mapped handle to a Vortex file. Implements AutoCloseable. Closing releases the mmap region;
all Array buffers obtained during scans become invalid.
| Method | Returns | Notes |
|---|---|---|
static open(Path) |
VortexReader |
Uses ReadRegistry.loadAll() |
static open(Path, ReadRegistry) |
VortexReader |
Custom registry (e.g. allowUnknown()) |
static open(Path, ReadRegistry, LayoutRegistry) |
VortexReader |
Custom layout decoders too |
dtype() |
DType |
Schema (typically DType.Struct) |
layout() |
Layout |
Layout tree (Struct → Zoned → Chunked → Flat) |
footer() |
Footer |
Segment specs, encoding specs |
version() |
int |
File format version |
fileSize() |
long |
File size in bytes |
scan(ScanOptions) |
ScanIterator |
Open a scan |
columnStats() |
Map<ColumnName, ArrayStats> |
Aggregated min/max per column |
slice(offset, length) |
MemorySegment |
Zero-copy slice of mmap region |
close() |
— | Releases mmap |
Handle to a remote Vortex file read via HTTP Range requests — no full-file download. Implements
VortexHandle. On open, fetches the last 65 KB (TAIL_SIZE) in one request to locate the trailer,
postscript, and metadata blobs, parsing Content-Range for the file size instead of a separate
HEAD round trip; each subsequent rawSegment(SegmentSpec) call during a scan fires one targeted
Range request. All fetched bytes share a single confined Arena, released on close().
| Method | Notes |
|---|---|
static open(URI) |
Uses ReadRegistry.loadAll(), a shared default HttpClient |
static open(URI, ReadRegistry) |
Custom decode registry |
static open(URI, ReadRegistry, HttpClient) |
Caller-supplied HttpClient (proxy, custom TLS, per-request timeout) |
static open(URI, ReadRegistry, LayoutRegistry, HttpClient) |
Custom decode registry, layout registry, and HttpClient |
close() |
Releases the confined Arena backing all fetched buffers |
The default HttpClient is a single JDK client shared across every VortexHttpReader instance
(never closed — its lifetime tracks the JVM); supply your own via the three-or-four-argument
open overloads to configure proxying, TLS, or timeouts per call site.
Writes a Vortex file. Implements Closeable. The file is complete and readable as soon as close() returns.
| Method | Notes |
|---|---|
static create(WritableByteChannel, DType.Struct, WriteOptions) |
Default codec set |
static create(WritableByteChannel, DType.Struct, WriteOptions, List<EncodingEncoder>) |
Custom encoder list; disables the global dict (would otherwise silently reach outside the given list). The list is also the whole cascade candidate set; as in Rust, a primitive no candidate accepts falls back to canonical vortex.primitive |
static create(WritableByteChannel, DType.Struct, WriteOptions, WriteRegistry) |
Custom WriteRegistry — encoders and extension encoders together |
writeChunk(Consumer<Chunk>) |
One batch of rows; typed builder validates column names + array types at each .put; missing columns throw IllegalStateException when the lambda returns. Preferred when columns are known at compile time. |
writeChunk(Map<ColumnName, Object>) |
One batch of rows by map. Validates that every schema column is present and that all columns share the same row count. Use when the column set is built dynamically (Parquet/JDBC importers, generic exporters). |
close() |
Finalizes file (footer, postscript, trailer) |
Batches are not chunks: as Rust's writer does, each primitive, boolean, decimal, string and binary
column is repartitioned into chunks of whole 8192-row blocks reaching 1 MB of uncompressed data
(e.g. 131 072 rows of I64, 524 288 of U16 dict codes), the rest becoming the last chunk. Nested
columns, and targets before core2026.08.0 (legacy vortex.stats zones), keep one chunk per batch.
Builder handed to the writeChunk(Consumer<Chunk>) lambda. Validates each .put
at the call site.
| Method | Notes |
|---|---|
put(ColumnName column, Object data) |
Adds one column; returns this for chaining |
Accepted array types per column DType:
DType |
Non-nullable array | Nullable array |
|---|---|---|
Primitive(I8/U8) |
byte[] |
Byte[] |
Primitive(I16/U16) |
short[] |
Short[] |
Primitive(I32/U32) |
int[] |
Integer[] |
Primitive(I64/U64) |
long[] |
Long[] |
Primitive(F32) |
float[] |
Float[] |
Primitive(F64) |
double[] |
Double[] |
Utf8 |
String[] |
String[] (nulls allowed) |
Bool |
boolean[] |
Boolean[] |
Decimal(precision, scale) |
BigDecimal[] |
BigDecimal[] (nulls allowed) |
Decimal values are rescaled to the column scale exactly: a value with more fractional digits than
the scale, or more digits than the precision, is rejected with IllegalArgumentException rather
than rounded. Every precision Rust supports (1–76, i8 … i256 storage) and negative scales work.
Record: (boolean enableZoneMaps, double compressionRatioThreshold, int allowedCascading, boolean globalDict, MemorySize globalDictMaxRetainedBytes, Map<EditionFamily, Edition> editions).
| Factory | Defaults |
|---|---|
WriteOptions.defaults() |
enableZoneMaps=true, compressionRatioThreshold=0.90, allowedCascading=3 (Rust's MAX_CASCADE), globalDict=true, compact=false, globalDictMaxRetainedBytes=MemorySize.ofGiB(2), editions={CORE: Editions.CORE_2026_08_3} |
WriteOptions.cascading(depth) |
Same defaults, allowedCascading=depth; 0 disables cascading (each column takes the first accepting encoder) |
| Method | Notes |
|---|---|
withZoneMaps(boolean) |
Toggle zone-map statistics. By default Rust's vortex.zoned: one zone per 8192 rows regardless of writeChunk batches, with max/min (strings: 64-byte bounds), nan_count for floats and null_count, no sum. An edition before core2026.08.0 gets the legacy vortex.stats: one zone per batch with a sum, pruned by vortex-jni only when every batch but the last has the same row count |
withGlobalDict(boolean) |
Toggle the shared cross-chunk dictionary |
withCompact(boolean) |
Rust's compact preset (with_compact): Zstandard for strings and binary and Pco for integers and floats compete in the cascade. Requires allowedCascading > 0; compact() reads it back |
withGlobalDictMaxRetainedBytes(long) |
Aggregate heap budget for buffered global-dict candidate columns |
withExecutor(Executor) |
Compress segments on executor (e.g. the common ForkJoinPool); default is the calling thread. Output is byte-identical to a sequential write. The writer owns arrays passed to writeChunk until close() — do not modify them |
withEdition(Edition) |
Enable an edition for its family, replacing any edition already enabled for that family |
withoutEditions() |
Turn the edition guard off, as Rust's disable_editions(): every encoding may be emitted, including those in no edition (fastlanes.delta, vortex.patched). Files may not be readable by other Vortex versions |
withColumnEncoding(ColumnName, ColumnEncoding) |
Choose one column's encodings from ColumnEncoding.candidates(EncodingId...) only — the whole candidate set, children included, with canonical encodings as the fallback (Rust's with_field_writer over a restricted scheme set). The edition guard still applies and the column skips the global dict. An unknown column, or an encoding the writer cannot emit, fails VortexWriter.create |
Record: (List<ColumnName> columns, RowFilter rowFilter, long limit) (built via columns(ColumnName...)). Empty columns = read all. NO_LIMIT =
Long.MAX_VALUE.
| Factory / builder | Effect |
|---|---|
ScanOptions.all() |
All columns, no filter, no limit |
ScanOptions.columns(ColumnName... names) |
Project columns |
ScanOptions.limit(long n) |
Limit rows |
.withColumns(ColumnName... names) |
Project columns (builder) |
.withFilter(RowFilter) |
Add zone-map filter |
.withLimit(long n) |
Cap rows |
.hasProjection() / .hasFilter() / .hasLimit() |
Predicates |
Sealed predicate tree used for zone-map pruning (per-chunk min/max) and row selection. Two
variants: RowFilter.Column(ColumnName, Predicate) binds one column to a value-test
Predicate (io.github.dfa1.vortex.reader.compute.Predicate — the same vocabulary the
compute kernels evaluate); RowFilter.And(List<RowFilter>) conjoins several. Chunks that
cannot match are skipped entirely.
| Static factory | Builds |
|---|---|
RowFilter.eq(col, val) |
Column bound to Predicate.Eq |
RowFilter.neq(col, val) |
Column bound to Predicate.Neq |
RowFilter.gt(col, val) |
Column bound to Predicate.Gt |
RowFilter.gte(col, val) |
Column bound to Predicate.Gte |
RowFilter.lt(col, val) |
Column bound to Predicate.Lt |
RowFilter.lte(col, val) |
Column bound to Predicate.Lte |
RowFilter.isNull(col) |
Column bound to Predicate.IsNull |
RowFilter.isNotNull(col) |
Column bound to Predicate.IsNotNull |
RowFilter.and(f1, f2, …) / f1.and(f2) |
And over the given filters |
Implements Iterator<Chunk> and AutoCloseable. Drives one scan.
| Method | Notes |
|---|---|
hasNext() |
Side-effect-free. Returns whether another chunk is available after zone-map pruning. |
next() |
Returns a fresh Chunk whose arena the caller closes. Throws IllegalStateException if a prior Chunk is still open, or NoSuchElementException if exhausted. |
forEachRemaining(Consumer) |
Overridden to wrap each next() in try-with-resources so chunks auto-close. |
close() |
Releases iterator state and closes any chunk still open. |
columnZoneStats(ColumnName) |
One ArrayStats per zone-map row; falls back to per-chunk stats when the column has no zone map. |
columnZones(ColumnName) |
One Zone(firstRow, rowCount, stats) per zone-map row, placed on the column's rows; empty when there is no usable zone map. |
Both return exact statistics only. Rust records string and binary zones as vortex.bounded_max/vortex.bounded_min, truncated bounds rather than exact values: filtered scans prune on them, but these methods leave min/max absent for such columns.
Implements AutoCloseable. Each chunk owns a confined Arena holding the decoded
columnar buffers; closing the chunk releases the arena. After close(), touching
any Array previously returned by column(...) or columns() raises FFM's scope
check (IllegalStateException).
Columns are stored as one order-preserving map keyed by the validated [ColumnName]; each
entry is a Chunk.Column(Array array, DType dtype) carrier, so a column's data and type can
never desync.
| Method | Notes |
|---|---|
rowCount() |
Rows in this chunk |
columns() |
SequencedMap<ColumnName, Chunk.Column>, schema order, unmodifiable |
<T extends Array> column(ColumnName name) |
Typed column lookup; throws VortexException if absent |
as(ColumnName name, Class<T> domainType) |
Extension column → typed List<T> |
isClosed() |
Whether close() has run |
close() |
Releases the chunk's arena. Idempotent. |
Immutable after construction. Build via ReadRegistry.builder() or the static convenience factories.
| Method | Notes |
|---|---|
static builder() |
Returns a fresh Builder |
static loadAll() |
Immutable registry populated with all built-in decoders |
static empty() |
Immutable empty registry (strict mode) |
hasDecoder(EncodingId) |
Lookup |
isAllowUnknown() |
Predicate |
| Method | Notes |
|---|---|
register(EncodingDecoder) |
Add a custom encoding decoder; throws if already registered |
registerDefaults() |
Add every built-in EncodingDecoder |
allowUnknown() |
Switch to passthrough mode — unknown nodes (and their children) decode as UnknownArray |
build() |
Produce the immutable ReadRegistry |
Register custom encoding decoders programmatically via register(EncodingDecoder) — there is no
ServiceLoader discovery. Extension decoders
(io.github.dfa1.vortex.reader.extension.ExtensionDecoder) are not registry-managed: the built-in
implementations are singletons invoked directly by their ExtensionId.
The write-side mirror of ReadRegistry: maps EncodingId → EncodingEncoder and ExtensionId →
ExtensionEncoder. Immutable after construction. Build via WriteRegistry.builder() or the static
convenience factories.
| Method | Notes |
|---|---|
static builder() |
Returns a fresh Builder |
static loadAll() |
Immutable registry populated with all built-in encoders + extensions |
static empty() |
Immutable empty registry |
encoderMap() |
Map<EncodingId, EncodingEncoder> for EncodeContext |
lookup(ExtensionId) |
Registered ExtensionEncoder for the id, or null |
| Method | Notes |
|---|---|
register(EncodingEncoder) |
Add a custom encoder; throws VortexException if its id is already registered |
register(ExtensionEncoder) |
Add an extension encoder, keyed by ExtensionEncoder#extensionId(); throws on duplicate |
registerDefaults() |
Add every built-in EncodingEncoder and ExtensionEncoder |
build() |
Produce the immutable WriteRegistry |
EncodingId is a sealed WellKnown/Custom type (see Identity types),
so register(EncodingEncoder) accepts a genuinely third-party encoding — encodingId() can return
new EncodingId.Custom("acme.myencoding"), and VortexWriter.create(channel, schema, options, registry) will pick it whenever its accepts(DType) matches. ExtensionId, by contrast, is a
closed enum with only the four spec-defined constants (VORTEX_DATE/VORTEX_TIME/
VORTEX_TIMESTAMP/VORTEX_UUID) — register(ExtensionEncoder) can only re-implement one of
those four, not introduce a wire-new extension type; adding a genuinely new extension type means
adding an ExtensionId constant in core itself (see CLAUDE.md § Adding an extension
type).
Frozen, named, additive sets of encoding IDs, each carrying a forever read-compatibility
guarantee once frozen (see ADR 0023). vortex-java mirrors Rust's vortex-edition declarations at
the pinned vortex-jni release (0.86.1); EditionCatalogParityIntegrationTest fails the build if the
two drift. Editions are a write/read policy, not part of the wire format: nothing about a
targeted edition is ever persisted into a .vortex file.
| Member | Notes |
|---|---|
CORE_2025_05_0 … CORE_2026_08_3 |
The frozen core editions (forever read-compatibility guarantee) |
PREVIEW_2026_08_0 |
The preview family's draft edition: opt-in components not yet in core (empty today) |
ZSTD_2026_02_0 |
The zstd family's draft edition (vortex.zstd_buffers), declared by Rust's vortex-zstd plugin |
ALL |
Every declared edition, in Rust's declaration order |
cumulativeMembers(Edition) |
The edition's own additions plus every earlier same-family edition's |
owningEdition(EncodingId) |
The edition an id first joined, or empty if it belongs to none |
vortex-java implements every encoding in the catalog; vortex.parquet.variant and zstd's
vortex.zstd_buffers are read only.
fastlanes.delta and vortex.patched belong to no edition, as in Rust.
WriteOptions.defaults()/cascading(depth) enable the newest frozen core edition
(Editions.CORE_2026_08_3), the edition Rust's default session enables, so a default write may emit
exactly what Rust's default writer may — including vortex.onpair, which competes with FSST for
string columns. If a write would emit an encoding outside the union of every enabled edition's
cumulative members, VortexWriter fails the write immediately with a VortexException naming the
encoding and the configured edition(s). Where possible (any selection routed through
CascadingCompressor, including nested competitions like a masked column's validity-bitmap cascade)
the guard instead steers selection away from the ineligible candidate ahead of time, falling back to
the best remaining one — see ADR 0023.
WriteOptions.withEdition(Edition) enables an edition, replacing any edition already enabled for
that edition's family; several families can be enabled at once, at most one edition per family.
WriteOptions.withoutEditions() turns the guard off, the counterpart of Rust's
disable_editions(): the cascade then also offers encodings that are in no edition
(fastlanes.delta), and a file written this way may not be readable by other Vortex versions.
When ReadRegistry hits an unregistered encoding id (and allowUnknown() is off), the thrown
VortexException names the edition the id belongs to, or says the id is unknown to every edition
and points at allowUnknown().
The vortex-fsst module is the standalone FSST (Fast Static Symbol Table) string-compression
algorithm, usable independently of Vortex. It depends only on the JDK (java.lang.foreign), never
on core/reader/writer; writer/reader depend on it. The vortex.fsst encoding adapter is
one caller — the module itself knows nothing of the Vortex wire format. See
ADR 0022. On the read side, FsstEncodingDecoder never
decompresses eagerly: it returns a LazyFsstVarBinArray that decompresses row i's code range only
when that row is actually read, since FSST's per-row code range is independent of every other row
(unlike Bitpacked/Pco/Zstd, which do need a decoded window) — see
ADR 0026.
Compress: train a Compressor over a corpus, then compress rows against its table. Decompress:
build a Decompressor (from the trained Compressor, or from raw table arrays). Hot-path methods
accept MemorySegment (FFM-native, matching the rest of the codebase); byte[] overloads exist for
zero-ceremony standalone use, and both paths produce identical output for the same input.
| Method | Notes |
|---|---|
new CompressorBuilder() |
Fresh builder using a fixed reproducible default seed |
seed(long) |
Override the training-sample seed; returns this for chaining |
train(byte[][] rows) |
Runs the bottom-up generation loop, returns a trained Compressor (deterministic in rows + seed) |
Immutable trained symbol table (up to 255 codes; 0xFF is the escape). Produced by
CompressorBuilder.train(byte[][]).
| Method | Notes |
|---|---|
symbolCount() |
Number of symbols in the table (0–255) |
packedSymbol(int code) |
The symbol's bytes packed LSB-first into a long |
symbolLength(int code) |
The symbol's length in bytes (1–8) |
codesSortedByLength() |
Code permutation, multi-byte length-ascending then all length-1 last |
toDecompressor() |
A Decompressor bound to this table |
compress(byte[] in, int start, int end, byte[] out, long outPos) |
Greedy longest-match compress; size out ≥ 2 * (end - start) |
compress(MemorySegment in, long start, long end, MemorySegment out, long outPos) |
FFM-native hot path; same sizing rule |
Decodes an FSST code stream against a trained table (parallel arrays indexed by code).
| Method | Notes |
|---|---|
static of(long[] packedSymbols, int[] lengths) |
Bind to a table; arrays are referenced, not copied, and must match in length |
static final int ESCAPE |
The escape code (0xFF) |
decompress(byte[] in, int start, int end, byte[] out, long outPos) |
Decode; size out ≥ 8 * (end - start) |
decompress(MemorySegment in, long start, long end, MemorySegment out, long outPos) |
Unconditional-8-byte-store decode; caller MUST allocate out with ≥ 7 bytes of trailing slack |
record Symbol(long packedBytes, int length) — up to 8 bytes packed LSB-first (byte k at bit
k * 8), length in 1–8.
| Method | Notes |
|---|---|
static of(byte[] data, int offset, int length) |
Pack length bytes at data[offset..] LSB-first |
byteAt(int i) |
The i-th byte (0-indexed; byte 0 is the LSB) |
Maps LayoutId → LayoutDecoder, making layout decode pluggable the way ReadRegistry makes
encodings pluggable. Programmatic registration only — no ServiceLoader. Unknown layouts fail
loudly (VortexException); there is no allow-unknown mode for layouts (Rust default).
| Method | Notes |
|---|---|
static defaults() |
The built-ins: flat, chunked, zoned (both aliases), dict, struct |
static builder() |
Returns a fresh Builder |
hasDecoder(LayoutId) |
Lookup |
decode(ctx, layout, dtype) |
Dispatch by the layout's typed id |
Builder: register(LayoutDecoder) (throws on duplicate id), registerDefaults(), build().
A decoder claims its ids via LayoutDecoder.layoutIds() (a set — the zoned decoder claims both
vortex.zoned and legacy vortex.stats). Custom decoders receive a LayoutDecodeContext
(decodeChild, decodeSegment access, arena) and are reachable end-to-end via
VortexReader.open(path, readRegistry, layoutRegistry).
Supports flat schemas and nested LIST/STRUCT columns, recursively composable (e.g.
LIST<STRUCT<...>>); MAP and VARIANT are not yet supported. Un-annotated BYTE_ARRAY maps
to DType.Binary.
| Method | Notes |
|---|---|
importParquet(Path in, Path out) |
Defaults |
importParquet(Path in, Path out, ImportOptions) |
Tuned |
importParquet(URI in, Path out) |
Remote source over HTTP(S), defaults |
importParquet(URI in, Path out, ImportOptions) |
Remote source over HTTP(S), tuned |
A URI source is fetched entirely through targeted HTTP Range requests (HttpInputFile,
internal) — no full-file download occurs.
Record: (int chunkSize, List<String> columns, ProgressListener progressListener, WriteOptions writeOptions).
| Factory / builder | Notes |
|---|---|
ImportOptions.defaults() |
chunkSize=65_536, no projection, WriteOptions.cascading(3) |
.withColumns(List<String>) |
Project columns during import |
.withProgressListener(listener) |
Progress callbacks |
.withWriteOptions(WriteOptions) |
Override write options |
.withChunkSize(int) |
Rows per writeChunk batch |
Supports flat schemas only: Bool, Primitive (F16 as Parquet's FLOAT16), Utf8, Binary, and the
vortex.timestamp extension over MILLIS/MICROS/NANOS resolution. Struct/List/Map top-level
columns throw UnsupportedOperationException — the inverse of ParquetImporter's nested
LIST/STRUCT support does not exist yet.
| Method | Notes |
|---|---|
exportParquet(Path in, Path out) |
Defaults |
exportParquet(Path in, Path out, ExportOptions) |
Tuned |
exportParquet(VortexHandle in, Path out) |
Already-open source (local VortexReader or remote VortexHttpReader), defaults |
exportParquet(VortexHandle in, Path out, ExportOptions) |
Already-open source, tuned |
The VortexHandle overloads don't close in — the caller opened it and keeps ownership of its
lifecycle, so a remote VortexHttpReader.open(uri) can be exported without an intervening local
copy.
Record: (List<String> columns, ProgressListener progressListener, WriterConfig writerConfig).
WriterConfig is Hardwood's own writer configuration type
(dev.hardwood.writer.WriterConfig) — row group/page targets, compression codec, column
encoding policy.
| Factory / builder | Notes |
|---|---|
ExportOptions.defaults() |
No projection, Hardwood's WriterConfig.defaults() |
.withColumns(List<ColumnName>) |
Project columns during export |
.withProgressListener(listener) |
Progress callbacks |
.withWriterConfig(WriterConfig) |
Override Hardwood writer configuration |
Column types are inferred from the data (long → double → boolean → utf8, in priority order)
unless a schema is given via ImportOptions#withSchema. Schema-driven CLI subcommands
(schema, count, …) are documented for Vortex/Parquet files only — CSV has no schema of its
own, so it is only ever an import source: never a destination (there is no CSV export from
import), and never a source for the other subcommands, which all expect a Vortex or Parquet
file.
| Method | Notes |
|---|---|
importCsv(Path in, Path out) |
Defaults |
importCsv(Path in, Path out, ImportOptions) |
Tuned |
importCsv(URI in, Path out) |
Remote source over HTTP(S), defaults |
importCsv(URI in, Path out, ImportOptions) |
Remote source over HTTP(S), tuned |
A URI source streams the response body directly, front to back, as it arrives — unlike
Parquet's random-access footer-first format, CSV needs no Range requests and no local temp file.
Reads rows from a JDBC ResultSet and writes a Vortex file. The schema is derived entirely from
ResultSetMetaData — no type inference, unlike CsvImporter. SQL NULL is stored as a zero/empty
placeholder plus a validity bit (0, 0.0, false, ""); vortex.date/vortex.time/
vortex.timestamp/vortex.uuid columns round-trip through the matching JDBC getter
(getDate/getTime/getTimestamp/getObject, the last handling java.util.UUID, byte[16], and
36-char string driver representations).
| Method | Notes |
|---|---|
static importTable(Connection, String tableName, Path out) |
SELECT * FROM tableName, default options |
static importQuery(Connection, String sql, Path out) |
Arbitrary SELECT, default options |
static importQuery(Connection, String sql, Path out, JdbcImportOptions) |
Arbitrary SELECT, tuned |
Record: (int fetchSize, int chunkSize, WriteOptions writeOptions, ProgressListener progressListener).
| Factory / builder | Notes |
|---|---|
JdbcImportOptions.defaults() |
fetchSize=10_000, chunkSize=65_536, WriteOptions.cascading(3), no listener |
.withFetchSize(int) |
Rows the JDBC driver fetches per round trip |
.withChunkSize(int) |
Rows per writeChunk batch; chunks on disk follow the writer's repartition |
.withWriteOptions(WriteOptions) |
Override write options |
.withProgressListener(listener) |
Callback invoked after each full chunk is flushed |
@FunctionalInterface: void onProgress(long rowsDone, long rowsTotal). rowsTotal is -1 —
JDBC has no cheap way to know the row count ahead of a SELECT COUNT(*).
Entry point for querying Vortex files with SQL through Apache
Calcite. connect(String, Map<String, Path>) folds the
DriverManager + unwrap + getRootSchema().add(...) boilerplate into one call, returning a plain
java.sql.Connection.
| Method | Notes |
|---|---|
static connect(String schemaName, Map<String, Path> tables) |
Opens a Calcite JDBC connection with tables registered under schemaName (schemaName.tableName in SQL); connection is the caller's to close |
The connection uses Calcite's JAVA lexical policy (case-sensitive identifiers, unquoted case
preserved) and the Babel SQL parser, so a column named after a reserved word — close, open,
value, year, … — is queryable unquoted (select close, open from vtx.ohlc). The exception is
keywords that open a typed literal (date, time, timestamp, interval), which stay reserved
even under Babel and need back-ticks: select `date` from vtx.ohlc.
Whole-table aggregates (min/max/sum) are answered from zone-map statistics via
VortexAggregatePushDownRule where possible — no data segment is decoded. The fold needs one zone
per chunk and a per-zone sum, i.e. the legacy vortex.stats zone map (writes targeting an edition
before core2026.08.0); files with Rust's vortex.zoned — the default since #447, and every
Rust-written file — answer the same aggregates through a scan. connect is the
supported entry point; VortexSchema/VortexTable (io.github.dfa1.vortex.calcite) are the
underlying schema/table plumbing it assembles, exposed for callers building a Calcite schema tree
by hand instead of going through connect.
The -all uber-jar is self-contained: pure-Java code plus the native libzstd for all
supported platforms (the FFM loader picks the right one at runtime) — no system libraries
required.
The cli module ships a fat jar with subcommands for inspecting and querying Vortex files.
./mvnw package -pl cli -am -DskipTests
java -jar cli/target/vortex-cli-*-all.jar <subcommand> [args]| Subcommand | Syntax | Description |
|---|---|---|
inspect |
inspect [--html] <file.vortex | http(s)://url> |
Layout tree, encodings, row counts, buffer sizes; --html emits a self-contained HTML report on stdout |
tui |
tui <file.vortex | http(s)://url> |
Interactive layout-tree browser (lazy stats + data) |
view |
view <file.vortex | http(s)://url> |
Interactive spreadsheet-grid browser over the row data |
schema |
schema <file.vortex> |
Column names and types |
count |
count <file.vortex> |
Total row count |
stats |
stats <file.vortex> |
Per-column min/max |
export |
export <file.vortex|url> [out.csv|out.parquet|-] |
All columns to CSV (default) or Parquet, by output extension; - for CSV on stdout. A url source requires an explicit out.parquet path — CSV/stdout from a URL isn't supported |
select |
select <file.vortex> <col> [col2 ...] [--where "<expr>" ...] |
Project columns to CSV; each --where is a filter expression and all must hold (the filtered columns need not be printed) |
filter |
filter <file.vortex> "<expr>" |
Filter rows to CSV |
import |
import [--delimiter <char>] [--compact] <file.csv|file.parquet|url> [out.vortex|out.parquet] |
CSV or Parquet (local or remote) source to Vortex; a .parquet output is CSV-only (chains through a temp Vortex file internally) — a Parquet source always produces Vortex, .parquet output is rejected; --compact is Rust's compact preset (Zstandard for text, Pco for numbers), see Convert CSV to Vortex |
Any subcommand also takes --timing, anywhere on the command line, including before the subcommand: it prints
elapsed: <n> ms to stderr when the command ends, so data on stdout stays clean in a pipeline. The time is printed
even if the command fails, and the exit status is unchanged. The clock starts when main starts, so it leaves out the
JVM's start-up (roughly half of a short command such as count: 33 ms of work in a 60 ms process on a 127 MB file).
For the interactive tui and view it includes the time you spend in them.
<column> <op> <value>
| Operator | Meaning |
|---|---|
>, >= |
Greater than, greater-or-equal |
<, <= |
Less than, less-or-equal |
=, == |
Equal |
!= |
Not equal |
Values are parsed as integer, double, boolean, or string (in that order). A string is written unquoted
(symbol = BTCUSDT): quote characters are not stripped, so symbol = 'BTCUSDT' compares against the text 'BTCUSDT'
(quotes included) and matches no row, without an error.
filter takes one comparison. For several conditions (all must hold) and to print only some columns, use select with
--where, repeated once per condition; an interval is two conditions:
select <file.vortex> open_time close --where "symbol = BTCUSDT" --where "open_time >= 1718452800000" --where "open_time < 1718456400000"
A --where column does not have to be among the printed columns. CSV output uses \r\n line endings.
8 bytes at EOF:
version (u16 LE) | postscriptLen (u16 LE) | magic ("VTXF")
The postscript is a FlatBuffer blob immediately before the trailer. It points (offset + length) to: the Footer (FlatBuffer), the DType (FlatBuffer), and the Layout (FlatBuffer) — each stored elsewhere in the file.
See explanation.md#memory-model for the mmap lifecycle.