The open
telemetry lake.
Siglake is an Apache-2.0, OTel-native telemetry store built on Parquet, Apache Iceberg, and DataFusion. Collect with OpenTelemetry, explore with SQL, and keep your data in your own S3 bucket as Parquet v2 in Apache Iceberg.
OpenTelemetry in. DataFusion SQL. Your bucket. Built in Rust.
# Point any OTel Collector at Siglake
exporters:
otlphttp:
endpoint: https://logs.example.com
headers:
X-Scope-OrgID: acme
otlp:
endpoint: siglake-ingester:4317
headers:
x-scope-orgid: acme
tls:
insecure: true
# …or POST /v1/logs directlyThe acme example requires a trusted gateway that sets this header and strips the client's value. Start the ingester with--trust-scope-header. Configure or disable OTLP/gRPC.
-- DataFusion SQL: joins, CTEs,
-- and window functions
SELECT host, count(*) AS n
FROM events
WHERE raw LIKE '%ERROR%'
GROUP BY host
ORDER BY n DESC# One connection tails this ingester's pre-WAL path
# Best-effort: no history or replay; through a load balancer,
# matching events on other ingesters are absent
curl -N -H 'X-Scope-OrgID: acme' \
https://logs.example.com/api/v1/stream
# …or cursor Iceberg to replay:
# the first poll reads the snapshot
# from --since, then each later poll
# delivers the rows added by new
# commits, late arrivals included
siglake subscribe --table events \
--since 2026-07-29T00:00:00ZThree things every modern telemetry store should be.
Your data, in your bucket, in open formats.
OTLP in the front, Parquet v2 out the back — on S3, cataloged by Apache Iceberg. Siglake-specific Puffin sidecars add optional indexes without replacing the Parquet data. And Siglake itself is Apache-2.0 open source.
- OTLP logs + traces · Elasticsearch-compatible bulk
- Parquet v2 + Iceberg, time-ordered and day-partitioned
- Trigram blooms and inverted indexes by default · optional Parquet-native blooms
Ingest your telemetry. No per-gigabyte license.
Siglake separates ingest, compaction, and query so you can scale each role independently. Your warehouse lives in your S3 bucket, with no Siglake license fee per gigabyte. You pay for the infrastructure you use.
- Separate ingest, compaction, and query roles
- One Iceberg warehouse, no hot/warm/cold tiers
- Durable WAL and SQL catalog alongside object storage
A stream your tools can hook into.
A live SSE tail at the front, Iceberg subscriptions underneath, SQL on top, and raw Parquet at the bottom. Your dashboards, agents, and batch tools can work with the exposed data and APIs.
- Live SSE tail: GET /api/v1/stream
- Incremental Iceberg subscriptions, late-arrival safe
- Elasticsearch-compatible _bulk · Jaeger trace-query subset
One open pipeline. Hook in at any stage.
OTLP in, an Arrow IPC write-ahead log (WAL), continuous commits into Iceberg on your object store, and SQL on top. Connect through ingest APIs, the Rust WAL consumer, Iceberg subscriptions, or SQL.
- INGEST
OTLP · logs · traces · ES bulk
↳ SSE live tail
ingest events, pre-WAL - WAL
Arrow IPC
acked on append - DRAIN
continuous Iceberg commits
- STORE
Iceberg · Parquet
your bucket · open formats - READ
SQL · Iceberg subscriptions
The live tail branches before the WAL append. SQL can also include uncommitted sealed WAL segments when --query-wal-buffer-dir is configured.
Tail this ingester’s events before the WAL append. Best-effort, with possible gaps and duplicates, no replay, and no guarantee that an event is accepted. Use Iceberg subscriptions for replay.
Start with a snapshot from --since, then follow rows added by new commits, late arrivals included.
Elasticsearch-compatible _bulk ingest, SQL search, and the Jaeger trace-query subset Grafana renders: services, operations, trace search and trace fetch.
Parquet v2 in Apache Iceberg, in your bucket. Read a committed snapshot using the catalog’s current metadata_location; see the example below.
Your warehouse is Parquet v2 in Apache Iceberg.
Every event ends up as a Parquet v2 file in your object store, cataloged by Apache Iceberg. Siglake-specific Puffin sidecars add optional indexes without replacing the Parquet data.
S3 · Iceberg · Parquet v2 · Puffin · Arrow · DataFusion
├── siglake/events/
│ ├── data/
│ │ └── day_ts=2026-09-05/ ← day partition
│ │ └── siglake-g1-019f…-00000.parquet ← Parquet v2
│ └── metadata/
│ ├── 00012-8c4e….metadata.json
│ ├── snap-7126…-0-3a91….avro ← manifests
│ └── siglake-agg/<table-uuid>/siglake-aggregates.json ← rollups
├── siglake/app_logs/ ← user index
└── siglake/traces_prod/ ← user indexThat reads the committed snapshot, not whatever files sit under the prefix. Take the path from the catalog's current metadata_location: Siglake writes no version-hint.text, so the bare table directory does not resolve. See how to find the current metadata file. Checked 2026-09-06 with DuckDB 1.5.5 on a local-disk warehouse.
Explore the evidence
Built for telemetry.
Measured in the open.
Open formats, independent ingest and query roles, and indexes designed for log search. Siglake brings these together so you can explore telemetry while keeping the underlying data in your own bucket.
Performance depends on the workload. Our benchmark site brings the measurements, resource assumptions, query coverage, and methodology together so you can judge the results for yourself.
Explore the benchmarksSee how query behavior varies across engines and workloads.
Explore measured resource use and the assumptions behind cost comparisons.
Read what was tested, how caches were handled, and where comparisons have limits.
Start streaming events into your own bucket with one command.
One script brings up the full stack: Postgres catalog, MinIO warehouse, ingester, compactor, and query server. Run scripts/up.sh. Production is a Helm chart (or the operator) and your S3 bucket.