Changelog
Release posts: Upgrade from 1.8.0 to 2.0.0 · What's new in 1.8.0 · What's new since 1.7.0 · What's new in 1.6.0
All notable changes to Beacon are documented here, newest first. Entries are grouped into Added (new features), Changed (behaviour or internal changes), and Fixed (bug fixes).
v2.0.0 — 2026-09-29
Beacon 2.0.0 is a major release. A 1.8.0 server does not upgrade in place. A 2.0.0 server does not read the table definitions of a 1.8.0 server, so create your tables again. Read Upgrade from 1.8.0 before you change the image tag. The full changelog lists each change.
Added
- N-dimensional execution. NetCDF, HDF5, Zarr, Atlas, BBF and GeoTIFF read through an nd pipeline. A
WHEREfilter and a projection run before the broadcast. See How it works. - A single-file database. One
beacon.dbfile holds the catalog, the managed tables and the users. See Storage internals. - File statistics. Beacon records the column ranges of each file. A query skips each file that holds no match. See File statistics.
- New readers. Read Iceberg tables and Icechunk repositories. Pure-Rust readers read netCDF and HDF5.
- 123 spatial functions with PostGIS names. See Spatial Functions.
- Secrets and remote catalogs.
CREATE SECRETholds credentials.ATTACHqueries another Beacon server. See ATTACH. - An embedded engine.
pip install beacondbruns the engine in your Python process.
Changed
- One product, one name. "Beacon Data Lake" and "BeaconDB" are now Beacon.
- The netCDF and HDF5 readers are pure Rust by default. Set
BEACON_NETCDF_USE_RUST_READER=falseorBEACON_HDF5_USE_RUST_READER=falseto use netCDF-C. - File statistics are on by default. Set
BEACON_FILE_STATS_ENABLE=falseto stop them. BEACON_S3_DATA_LAKEis nowBEACON_S3_DATASETS. The old name still works.- The minimum Rust version is 1.94. A build from source also needs PROJ 9.6.2 or later.
Removed
st_within_pointandst_geojson_as_wkt. Use the PostGIS functions, such asST_Within.
v1.8.0 — 2026-06-27
Added
- Authentication & role-based access control. Beacon gains a built-in auth/RBAC layer: a single config-defined super-user (
BEACON_ADMIN_*) that owns all writes and management, read-only SQL-managed users and roles (CREATE USER/CREATE ROLE/GRANT/DENY/REVOKE), andSELECTgrants/denies scoped to tables (ON TABLE) or file globs (ON PATH) with deny-wins, default-deny evaluation. Enforcement is opt-in (BEACON_AUTH_ENFORCE), anonymous access is configurable (BEACON_AUTH_ANONYMOUS_ENABLED), and OIDC bearer tokens are supported (BEACON_OIDC_*) with Beacon still owning the grant model. See the Access Control guide. - Lance-backed managed tables. Managed tables (
CREATE TABLE) are now backed by Lance: fast local CRUD, native updates/deletes via deletion vectors, and secondaryINDEXes (btree, bitmap, full-text). Lance is the sole managed-table engine — the short-lived Apache Iceberg managed engine has been removed (Iceberg will return as an external, read-in-place source). See CREATE TABLE (Managed). - Delta Lake tables. Read Delta Lake tables in place with
read_delta()orCREATE EXTERNAL TABLE … STORED AS DELTA, including time travel, and append withINSERT INTO. See the Delta Lake chapter. - PostgreSQL / MySQL external tables.
CREATE EXTERNAL TABLE … STORED AS POSTGRES | MYSQLregisters a federated pointer at a table in an external relational database; Beacon pushes filters, projected columns,LIMIT, and aggregates down to the source. Passwords are encrypted at rest withBEACON_SECRETS_KEY. See the SQL Databases chapter. - Crawlers.
CREATE CRAWLERdiscovers files in the datasets store and registers external tables automatically — AWS Glue-style — with partition detection and triggers, plus admin REST endpoints to list, run, and delete them. See the Crawlers chapter. - Admin Web UI. A React admin console is now bundled into the Beacon server and Docker image and served at
/admin: an Athena-style SQL workbench plus pages to manage tables, datasets, crawlers, and external tables. See the Admin Web UI guide. - TypeScript / JavaScript SDK (
@beacon/client). An isomorphic SDK for Node.js and the browser, with an EF Core / LINQ-style query builder and Arrow result decoding. See the TypeScript SDK guide. beacon-datalake-cli. A Python terminal client (interactive REPL and one-shot subcommands) that runs SQL, explores tables / datasets / schemas, and exports to CSV, Parquet, Arrow IPC, or NetCDF. See the Beacon Datalake CLI guide.EXPLAIN ANALYZEendpoint.POST /api/explain-analyze-queryruns a query and returns the physical plan annotated with per-operator runtime metrics (rows, bytes, time). See the API querying chapter.- Per-format configuration & per-table
OPTIONS. NetCDF, Atlas, and BBF accept per-format configuration and per-tableOPTIONS (...)clauses. - Broadcast-compatible
SELECT *for n-dimensional data.SELECT *over NetCDF and Zarr datasets auto-narrows to a broadcast-compatible default set of dimensions, with the selection now robust across irregular variable shapes.
Changed
- Runtime-owned configuration. Configuration is threaded through
Runtime::new(Arc<Config>)instead of being process-global. The storage config owns the data directories and optional S3 settings, withS3Configas the single source of truth for object storage. - Slimmer data lake. The bespoke
TableManager/FileManager/DataLakelayers were replaced with a native DataFusionSessionContext. table-configis now an admin-only REST endpoint (GET /api/admin/table-config).- Legacy logical tables are read as external tables, so tables created by older versions keep resolving.
- Reduced dependencies and tuned the build for faster compile times.
- File-format table functions accept a single path/glob string in addition to a list, and the dataset-schema endpoint now accepts
.tifas well as.tiff. - Enriched OpenAPI / Swagger documentation for the REST API.
Fixed
- A batch of bugs surfaced by an expanded integration-test suite: CTAS /
INSERTrow truncation, Lance string-predicateUPDATE/DELETE,ALTER TABLE ADD COLUMNforVARCHAR/Utf8View, inverted-index lookups, MySQL TLS connections, DeltaINSERTreopen, n-dimensionalcount(*)returning0, PostgreSQL / MySQL federated execution, NetCDF explicit-dimension subsetting, andread_zarrdirectory listing. EXPLAIN ANALYZEpanic over external NetCDF / n-dimensional tables.- Root redirect now always points to the Swagger UI.
- Clear error for output formats on non-
SELECTstatements. Requesting an output format for a statement that produces no result set (DDL/DML) now returns an explanatory error instead of failing obscurely. super_typeArrow coercion bugs surfaced by expanded unit-test coverage.
v1.7.3 — 2026-06-19
Added
- GeoParquet read support.
.geoparquetfiles are read with their geometry columns decoded to native GeoArrow; query them with theread_geoparquet()table function or register aSTORED AS GEOPARQUETexternal table. See the GeoParquet SQL chapter and the data-lake GeoParquet chapter. - Federated remote tables.
CREATE EXTERNAL TABLE … STORED AS REMOTEpoints at a table on another Beacon instance over Arrow Flight SQL, pushing filters, projected columns,LIMIT, and whole joins/aggregates down to the remote so only the reduced result set crosses the network. See the remote-tables SQL chapter and the federation setup chapter.
v1.7.2 — 2026-06-16
Changed
- Error handling & logging overhaul. The storage and format crates — NetCDF, Zarr, GeoTIFF, Atlas, object storage, configuration, the DataFusion extensions, and the n-dimensional array layers — now surface structured, contextual errors and clearer log output on read and configuration failures.
Fixed
- Docker image build. Corrected the
Dockerfileso the image builds again after the query-engine crate consolidation (it referenced crates that no longer exist).
v1.7.1 — 2026-06-15
Changed
- Unified query execution. JSON DSL and SQL queries now run through a single physical execution pipeline in
beacon-core, driven by one custom physical planner for all statements. The standalonebeacon-queryandbeacon-plannercrates were removed. - Simplified data lake. All table types are now described by a single
LogicalTableDefinitionmodel, replacing the previous separate file-CRUD path.
Fixed
- Swagger UI / Scalar redirects now resolve correctly when Beacon is served under a non-empty base path (
BEACON_BASE_PATH).
v1.7.0 — 2026-06-10
Added
- Row mutations on managed tables. Alongside
INSERT, managed tables now supportDELETE ... WHERE,UPDATE ... SET ... WHERE, andCREATE TABLE AS SELECT.DELETEandUPDATEare copy-on-write. - Schema evolution with
ALTER TABLEon managed tables —ADD COLUMN,DROP COLUMN,RENAME COLUMN, andALTER COLUMN ... TYPE(safe widening promotions). Existing rows keep reading correctly: added columns readNULLand renames preserve values. See CREATE TABLE (Managed). - CF
calendarsupport. CF time-unit parsing now honours the optional CFcalendarattribute, so non-Gregorian calendars are interpreted correctly. - SeaDataNet L05 mappings. New UDFs map SeaDataNet instrument L05 codes for salinity and temperature.
Changed
- Managed tables are now backed by Apache Iceberg instead of the previous Parquet-manifest format, giving them an ACID, schema-tracked, snapshot-based storage layer. Data and metadata live in Beacon's internal area of the configured storage (local or S3), alongside the datasets.
- CF time parsing for NetCDF and Zarr is centralized in one module, replacing the previous per-backend regex parsing.
- Global NetCDF attributes are now surfaced with a leading dot (e.g.
.Conventions) to cleanly distinguish them from variable attributes. - Filesystem event watching (
BEACON_ENABLE_FS_EVENTS) now defaults to enabled, so new files in watched datasets are picked up automatically. - Removed periodic table auto-sync and its
BEACON_TABLE_SYNC_INTERVAL_SECSsetting — managed tables are now transactional and no longer need background refreshes.
v1.6.1 — 2026-06-04
Added
- Atlas file format. Atlas is a directory-based array store — a single
atlas.jsonregistry describing one or more datasets — that Beacon discovers and queries automatically, just like Parquet or Zarr. Query a store with theread_atlas()table function or register it as an external table withSTORED AS ATLAS. Atlas keeps per-dataset column statistics, so Beacon prunes whole datasets before reading any array data and only loads the projected arrays. This makes re-encoding large NetCDF or Zarr collections into a single Atlas collection the recommended way to speed up repeated spatial/temporal range queries. See github.com/maris-development/atlas. - Materialized views.
CREATE MATERIALIZED VIEWruns a query once and persists the result as Parquet, so repeated, aggregation-heavy queries read straight from the cached result instead of recomputing. The newREFRESHstatement does a full recompute and atomically swaps in the new result, leaving the previous result intact if the refresh fails. - Configurable base path via the
BEACON_BASE_PATHenvironment variable. The HTTP API, OpenAPI document, and Swagger UI can be served under a path prefix, making it easier to run Beacon behind a reverse proxy or on a shared subpath. - LZW-compressed, stripped GeoTIFFs can now be decompressed, broadening the range of GeoTIFF/COG files Beacon can read.
- Zarr v3 support. Zarr reading moved onto Beacon's shared n-dimensional array engine, adding Zarr v3 support and predicate pushdown for Zarr-backed datasets.
Changed
- Upgraded the query engine to DataFusion 53 and Arrow 58. The Arrow version is now configurable at build time to ease integration with downstream tooling.
Fixed
- EDMO code extraction now uses the last set of parentheses in a SeaDataNet originator string, so institution names that themselves contain parentheses are mapped to the correct EDMO code.
v1.6.0 — 2026-05-08
Added
- Flight SQL. In addition to the existing HTTP query endpoint, datasets can be queried over the Flight SQL protocol — a more efficient, Arrow-native interface that opens Beacon up to clients such as JetBrains DataGrip, DBeaver, and other Apache Arrow Flight and BI tools.
- SQL tables. Create custom tables backed by Parquet files (local or in the cloud), populated from other files or bare
INSERTstatements — not tied to an existing dataset. - SQL views. Define views on top of datasets and expose them via the API, combining and transforming underlying datasets into purpose-built shapes.
- GeoTIFF support. Read and query TIFF files, including GeoTIFF and Cloud-Optimized GeoTIFF (COG), via the new
read_tiff()table function — adding raster data alongside Beacon's tabular formats. - ODV ASCII support. ODV ASCII files can be read directly as datasets and queried via the API, and streamed over the S3 protocol for efficient access to large ODV datasets in S3-compatible object storage.
- Ragged NetCDF & Zarr datasets. Read datasets whose variables have differing lengths and don't conform to a rectangular structure.
- NetCDF chunked streaming. Large NetCDF files are read in chunks rather than loaded wholesale into memory, reducing memory use and improving performance.
- NetCDF coordinate filter pushdown. Filters on coordinate variables are applied during reading, so far less data is read and processed for bounded spatial/temporal queries.
- NetCDF statistics & partition pruning. Per-column min/max statistics let Beacon prune whole files before reading, with new metadata table functions (
view_dataset_statistics,view_external_table_statistics,view_statistics_cache) to inspect cached statistics. - Merged tables. A table type that references and combines data from other tables, with dependency tracking that prevents deleting a table a merged table still relies on.
Fixed
- NetCDF scalar attributes. Attributes stored as a single-element list were not read correctly; all scalar attributes are now read properly.
v1.5.4 — 2026-01-05
Added
- SQL querying. Query datasets using SQL syntax in addition to the existing JSON query format.
- SQL querying for native datasets (e.g. NetCDF, Parquet) directly, without first creating a collection.
- ODV output units. Set the
unitfield per column in the ODV output schema, improving the clarity of the output data.
Fixed
- NetCDF FillValue is now set correctly on output, so missing values are properly represented.
Docs
- Added querying examples for both SQL and JSON query formats.
v1.2.0 — 2025-09-01
Added
- S3 / object storage. All file sources are abstracted via the Object Store crate, supporting backends such as MinIO, AWS S3, and Cloudflare R2.
- ODV ASCII streaming over the S3 protocol.
- NetCDF cloud reading using
#mode=bytes(http/https cloud storage endpoints only). - New table types: Preset Tables (data collections with metadata descriptions) and Geo-Spatial Tables (geo-spatial collections with metadata descriptions).
Changed
- Rewrote file-source reading to support S3.
- Updated dependencies for the latest Arrow, DataFusion, and GeoParquet.
v1.0.1 — 2025-05-05
Added
- SQL querying (#6) and SQL querying for native datasets such as NetCDF and Parquet (#33).
- ODV output units. Set the
unitfield per column in the ODV output schema (#52).
Fixed
- NetCDF FillValue / missing-value flag is now set correctly on output (#45).
Docs
- Added querying examples for both SQL and JSON query formats.