Skip to content

Getting Started ​

This guide deploys a Beacon node with Docker. The Quick Start below is the fastest path: start the server, add data, then explore it in the admin UI. The Local and S3 sections show production-shaped Docker Compose setups. The beacon-example repository holds ready-made examples with MinIO and sample datasets.

Running a node is four jobs, in this order:

  1. Deploy it, on this page.
  2. Configure the ports, the datasets store and the resource limits.
  3. Register your data as tables and views, so clients query names instead of paths.
  4. Secure and expose it, then point clients at it.

Prerequisites ​

Quick Start ​

Start the server and run your first query in a few minutes.

1. Run Beacon ​

Open the folder that holds your data. Then run:

bash
docker run -d \
  --name beacon \
  -p 5001:5001 \
  -e BEACON_ADMIN_USERNAME=admin \
  -e BEACON_ADMIN_PASSWORD=securepassword \
  -v ./datasets:/beacon/data/datasets \
  -v ./tables:/beacon/data/tables \
  ghcr.io/maris-development/beacon:v2.0.0

The tag v2.0.0 is the release that this documentation describes.

Beacon now serves on http://localhost:5001. That page is the home page. It shows the server version. It links to the admin UI, the Swagger UI, the API reference, the OpenAPI document, the health check and this documentation.

2. Add data ​

Copy supported files into the ./datasets folder. Supported files include .parquet, .nc, .zarr and .csv. Beacon finds them automatically. There is no import step.

3. Explore in the Admin UI ​

Open http://localhost:5001/admin. Sign in with the admin user name and password from step 1 (admin / securepassword). The server and the Docker image include the admin web UI. You deploy nothing extra. The UI gives you:

  • Query editor: write SQL, run it (⌘/Ctrl + Enter), read the results and download CSV or Parquet.
  • Datasets: browse the files that Beacon found and inspect their schemas.
  • Tables: create and manage tables over your datasets.
  • Crawlers and external tables: automate discovery and register external sources.
  • Server: runtime information, health and the available functions.

4. Or query over HTTP ​

Every request goes to one endpoint. Beacon streams back a file in the format that you ask for:

bash
curl -X POST http://localhost:5001/api/query \
  -H "Content-Type: application/json" \
  -d '{
    "sql": "SELECT * FROM read_parquet([\"**/*.parquet\"]) LIMIT 10",
    "output": { "format": "csv" }
  }'

The interactive API docs are at http://localhost:5001/swagger/.

Local ​

Write a docker-compose.yml for a repeatable setup. The longer docker run command below does the same. Point the volume paths at your datasets:

bash
docker run -d \
    --name beacon \
    --restart unless-stopped \
    -p 5001:5001 \
    -p 32011:32011 \
    -e BEACON_ADMIN_USERNAME=admin \
    -e BEACON_ADMIN_PASSWORD=securepassword \
    -v ./datasets:/beacon/data/datasets \
    -v ./tables:/beacon/data/tables \
    -v ./logs:/beacon/logs \
    ghcr.io/maris-development/beacon:v2.0.0
yaml
services:
    beacon:
        image: ghcr.io/maris-development/beacon:v2.0.0
        container_name: beacon
        restart: unless-stopped
        ports:
            - "5001:5001"   # HTTP API
            - "32011:32011" # Arrow Flight SQL
        environment:
            - BEACON_ADMIN_USERNAME=admin
            - BEACON_ADMIN_PASSWORD=securepassword
        volumes:
            - ./datasets:/beacon/data/datasets
            - ./tables:/beacon/data/tables
            - ./logs:/beacon/logs

For Compose, run docker compose up -d. Beacon now runs. Open http://localhost:5001 for the home page. It links to the admin UI and the API docs. Open the admin UI at http://localhost:5001/admin to explore and query. Open http://localhost:5001/swagger for the API docs. You can query any file in ./datasets at once.

Log files

The ./logs volume writes the log files to your machine. Beacon starts one file each day, for example beacon.log.2026-08-19. Without the volume the files stay in the container. See Log files.

Two ways to connect

Beacon exposes two endpoints. The HTTP API on port 5001 serves the home page, SQL and JSON queries, the admin UI and the OpenAPI docs. The Arrow Flight SQL server on port 32011 uses a columnar protocol with high throughput. Clients such as JetBrains DataGrip and the Python ADBC driver use it. Flight SQL authenticates with a bearer token. Tune it or switch it off with the BEACON_FLIGHT_SQL_*settings.

Secure your instance

The BEACON_ADMIN_* credentials protect the admin UI and all write operations. Change them from the defaults before you expose Beacon. To control who reads data, switch on access control with BEACON_AUTH_ENFORCE=true.

S3-Compatible Object Storage ​

Add the S3 environment variables. Then remove the datasets volume:

bash
docker run -d \
    --name beacon \
    --restart unless-stopped \
    -p 5001:5001 \
    -p 32011:32011 \
    -e BEACON_ADMIN_USERNAME=admin \
    -e BEACON_ADMIN_PASSWORD=securepassword \
    -e AWS_ENDPOINT=https://s3.amazonaws.com \
    -e AWS_ACCESS_KEY_ID=your-access-key \
    -e AWS_SECRET_ACCESS_KEY=your-secret-key \
    -e BEACON_S3_BUCKET=your-bucket-name \
    -e BEACON_S3_DATASETS=true \
    -v ./tables:/beacon/data/tables \
    -v ./logs:/beacon/logs \
    ghcr.io/maris-development/beacon:v2.0.0
yaml
services:
    beacon:
        image: ghcr.io/maris-development/beacon:v2.0.0
        container_name: beacon
        restart: unless-stopped
        ports:
            - "5001:5001"
            - "32011:32011"
        environment:
            - BEACON_ADMIN_USERNAME=admin
            - BEACON_ADMIN_PASSWORD=securepassword
            - AWS_ENDPOINT=https://s3.amazonaws.com
            - AWS_ACCESS_KEY_ID=your-access-key
            - AWS_SECRET_ACCESS_KEY=your-secret-key
            - BEACON_S3_BUCKET=your-bucket-name
            - BEACON_S3_DATASETS=true
        volumes:
            - ./tables:/beacon/data/tables
            - ./logs:/beacon/logs

Anonymous / public buckets

For a public bucket, remove AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY. Add AWS_SKIP_SIGNATURE=true instead.

For Compose, run docker compose up -d. You can query the files in the S3 bucket at once. The ./tables volume keeps the external tables and views that you create.

Next steps ​

Configure the node

Every settingConfiguration
Put the datasets on a bucketObject Storage
Memory, concurrency, cachesPerformance Tuning

Register the data

Name a set of filesExternal Tables
Save a queryViews · Materialized Views
Register a large tree on a scheduleCrawlers
Own the rows yourselfManaged Tables
Reach another node or a databaseATTACH · SQL Databases

Expose it

Decide who reads whatAccess Control
Explore in the browserAdmin Web UI
Point clients at itPython · TypeScript · CLI · DataGrip · Python ADBC
Document the query APIREST API

When something is wrong

Troubleshooting · FAQ

Released under the AGPL-3.0 License.