The open
telemetry lake.

Siglake is an Apache-2.0, OTel-native telemetry store built on Parquet, Apache Iceberg, and DataFusion. Collect with OpenTelemetry, explore with SQL, and keep your data in your own S3 bucket as Parquet v2 in Apache Iceberg.

OpenTelemetry in. DataFusion SQL. Your bucket. Built in Rust.

Ingest · query · tail
1. Ingest (OTLP)
# Point any OTel Collector at Siglake
exporters:
  otlphttp:
    endpoint: https://logs.example.com
    headers:
      X-Scope-OrgID: acme
  otlp:
    endpoint: siglake-ingester:4317
    headers:
      x-scope-orgid: acme
    tls:
      insecure: true

# …or POST /v1/logs directly

The acme example requires a trusted gateway that sets this header and strips the client's value. Start the ingester with--trust-scope-header. Configure or disable OTLP/gRPC.

2. Query (SQL)
-- DataFusion SQL: joins, CTEs,
-- and window functions
SELECT host, count(*) AS n
FROM events
WHERE raw LIKE '%ERROR%'
GROUP BY host
ORDER BY n DESC
3. Tail (any tool)
# One connection tails this ingester's pre-WAL path
# Best-effort: no history or replay; through a load balancer,
# matching events on other ingesters are absent
curl -N -H 'X-Scope-OrgID: acme' \
  https://logs.example.com/api/v1/stream

# …or cursor Iceberg to replay:
# the first poll reads the snapshot
# from --since, then each later poll
# delivers the rows added by new
# commits, late arrivals included
siglake subscribe --table events \
  --since 2026-07-29T00:00:00Z

Three things every modern telemetry store should be.

Open by design

Your data, in your bucket, in open formats.

OTLP in the front, Parquet v2 out the back — on S3, cataloged by Apache Iceberg. Siglake-specific Puffin sidecars add optional indexes without replacing the Parquet data. And Siglake itself is Apache-2.0 open source.

  • OTLP logs + traces · Elasticsearch-compatible bulk
  • Parquet v2 + Iceberg, time-ordered and day-partitioned
  • Trigram blooms and inverted indexes by default · optional Parquet-native blooms
Stream-scale economics

Ingest your telemetry. No per-gigabyte license.

Siglake separates ingest, compaction, and query so you can scale each role independently. Your warehouse lives in your S3 bucket, with no Siglake license fee per gigabyte. You pay for the infrastructure you use.

  • Separate ingest, compaction, and query roles
  • One Iceberg warehouse, no hot/warm/cold tiers
  • Durable WAL and SQL catalog alongside object storage
An open stream

A stream your tools can hook into.

A live SSE tail at the front, Iceberg subscriptions underneath, SQL on top, and raw Parquet at the bottom. Your dashboards, agents, and batch tools can work with the exposed data and APIs.

  • Live SSE tail: GET /api/v1/stream
  • Incremental Iceberg subscriptions, late-arrival safe
  • Elasticsearch-compatible _bulk · Jaeger trace-query subset
The open stream

One open pipeline. Hook in at any stage.

OTLP in, an Arrow IPC write-ahead log (WAL), continuous commits into Iceberg on your object store, and SQL on top. Connect through ingest APIs, the Rust WAL consumer, Iceberg subscriptions, or SQL.

  1. INGEST

    OTLP · logs · traces · ES bulk

    ↳ SSE live tail
    ingest events, pre-WAL

  2. WAL

    Arrow IPC
    acked on append

  3. DRAIN

    continuous Iceberg commits

  4. STORE

    Iceberg · Parquet
    your bucket · open formats

  5. READ

    SQL · Iceberg subscriptions

The live tail branches before the WAL append. SQL can also include uncommitted sealed WAL segments when --query-wal-buffer-dir is configured.

Live tail

Tail this ingester’s events before the WAL append. Best-effort, with possible gaps and duplicates, no replay, and no guarantee that an event is accepted. Use Iceberg subscriptions for replay.

GET /api/v1/stream
Iceberg subscription

Start with a snapshot from --since, then follow rows added by new commits, late arrivals included.

siglake subscribe --table events
Compatible APIs

Elasticsearch-compatible _bulk ingest, SQL search, and the Jaeger trace-query subset Grafana renders: services, operations, trace search and trace fetch.

_bulk · /api/v1/sql · jaeger/{index}/api
The lake itself

Parquet v2 in Apache Iceberg, in your bucket. Read a committed snapshot using the catalog’s current metadata_location; see the example below.

iceberg_scan('…/siglake/events/metadata/00012-….metadata.json')
Open data

Your warehouse is Parquet v2 in Apache Iceberg.

Every event ends up as a Parquet v2 file in your object store, cataloged by Apache Iceberg. Siglake-specific Puffin sidecars add optional indexes without replacing the Parquet data.

S3 · Iceberg · Parquet v2 · Puffin · Arrow · DataFusion

warehouse/s3://your-bucket
├── siglake/events/
│   ├── data/
│   │   └── day_ts=2026-09-05/                  ← day partition
│   │       └── siglake-g1-019f…-00000.parquet  ← Parquet v2
│   └── metadata/
│       ├── 00012-8c4e….metadata.json
│       ├── snap-7126…-0-3a91….avro             ← manifests
│       └── siglake-agg/<table-uuid>/siglake-aggregates.json  ← rollups
├── siglake/app_logs/       ← user index
└── siglake/traces_prod/    ← user index
$ duckdb -c "SELECT * FROM iceberg_scan('…/siglake/events/metadata/00012-8c4e….metadata.json') LIMIT 10"

That reads the committed snapshot, not whatever files sit under the prefix. Take the path from the catalog's current metadata_location: Siglake writes no version-hint.text, so the bare table directory does not resolve. See how to find the current metadata file. Checked 2026-09-06 with DuckDB 1.5.5 on a local-disk warehouse.

Explore the evidence

Built for telemetry.
Measured in the open.

Open formats, independent ingest and query roles, and indexes designed for log search. Siglake brings these together so you can explore telemetry while keeping the underlying data in your own bucket.

Performance depends on the workload. Our benchmark site brings the measurements, resource assumptions, query coverage, and methodology together so you can judge the results for yourself.

Explore the benchmarks
01 / LATENCY

See how query behavior varies across engines and workloads.

02 / RESOURCES

Explore measured resource use and the assumptions behind cost comparisons.

03 / METHODOLOGY

Read what was tested, how caches were handled, and where comparisons have limits.

Read the methodology ↗

Start streaming events into your own bucket with one command.

One script brings up the full stack: Postgres catalog, MinIO warehouse, ingester, compactor, and query server. Run scripts/up.sh. Production is a Helm chart (or the operator) and your S3 bucket.

$ git clone https://github.com/siglake/siglake && cd siglake
$ scripts/up.sh