# Iceberg: implementing the agent-queryable service Handoff spec for whoever builds the Iceberg layer. The Postgres equivalent is already shipped (`infra/ops/dbq.sh`) — **mirror it**. This document is the contract, the environment facts you'll trip over, and how the result gets verified. **Historical note:** `infra/ops/iceberg.sh` now exists, originally built against the environment facts below as they stood before the vps1/vps2 split — `beta` at `75.119.132.91` and `prod` at `169.58.46.181`, each its own box. `iceberg.sh` (and `dbq.sh`) have since been updated to resolve their SSH target and host facts from `deploy/env-topology.sh` instead of a hardcoded table, so they already reflect the current split (vps1 = `75.119.132.91` = dev + Forgejo + Grafana; vps2 = `169.58.46.181` = prod AND beta). This document is left as a historical record of the original design contract, not a live description of where each environment's engine runs — see `infra/ops/README.md` for the current behavior. ## Definition of done `infra/ops/iceberg.sh` exists and behaves like `dbq.sh`: ```sh infra/ops/iceberg.sh prod -c "select count(*) from ." infra/ops/iceberg.sh dev -c "show tables" echo "select 1" | infra/ops/iceberg.sh beta -f - ``` An agent (or operator) can query Iceberg in any environment, read-only, with no new network exposure and no interactive steps. ## The contract (non-negotiable) 1. **Same interface shape as `dbq.sh`** — `iceberg.sh [flags] ""`. Extra flags pass through to the engine; **stdin is forwarded** so `-f -` works. Non-interactive, deterministic, safe to call from CI. 2. **Read-only enforced by the system, not by convention.** A dedicated read-only identity — not an admin/superuser credential. You must be able to *demonstrate* a refused write (see Acceptance). 3. **No new network endpoints.** Run the engine where the data is already reachable (containerised), rather than publishing ports or opening firewall rules. See the prod constraint below — this is not negotiable, it's physics. 4. **Env-keyed dispatch, names resolved at call time.** Prod is Docker Swarm; task container names change on **every redeploy**. Resolve via `docker ps --filter name=…`; never hardcode a container name. 5. **No secrets in the repo** or in workflow files. ## Environment facts you need **As originally written** (pre vps1/vps2 split; kept for history — see the note at the top of this document): - Monorepo `Jinemi/thermograph` (domain dirs: `backend/ frontend/ infra/ observability/`). Ops tooling lives in `infra/ops/`. - Environments: **dev** = LAN compose stack on the dev machine; **beta** = `75.119.132.91`; **prod** = `169.58.46.181`. SSH as `agent` with `~/.ssh/thermograph_agent_ed25519` (passwordless sudo on both boxes). - **Prod runs Docker Swarm.** Its app database sits on the overlay network `thermograph_internal` (`10.0.2.0/24`), which **the prod host itself cannot route to**. Consequence, learned the hard way: `ssh -L` tunnelling works for beta but is *impossible* for prod. Any design that depends on a tunnel or a published port will fail on prod — that's why `dbq.sh` execs into the container instead. Assume the same for anything you deploy there. - The overlay is `attachable=true`, and the **dev machine is a Swarm worker** in the prod cluster (NodeAddr `10.10.0.3`) — so `docker run --network thermograph_internal …` from the dev box is a plausible route to prod-side services. Untested; verify before relying on it. - WireGuard mesh: prod `10.10.0.1`, beta `10.10.0.2`, dev `10.10.0.3`. - `rclone` and `age` are already installed on prod and beta. **Current facts, post vps1/vps2 split** (the boxes and mesh IPs above did not move — only which environment runs where, and what the boxes are called): - **vps2** = `169.58.46.181` = mesh `10.10.0.1` = the box called "prod" above. Runs **both** prod and beta now, as two separate Docker Swarm stacks — beta is a Swarm service too today, not a compose exception, so the tunnelling constraint above applies equally to it: `ssh -L` is impossible for beta now as well, not just prod. - **vps1** = `75.119.132.91` = mesh `10.10.0.2` = the box called "beta" above. Runs Forgejo, Grafana/Loki/Alloy, and **dev** (its own Postgres container, reached over SSH like any fleet host — not the LAN-box-local case the original facts describe). - **desktop** = mesh `10.10.0.3` = "the dev machine" above. Hosts no Thermograph environment any more; it's a Swarm worker for flex capacity and AI-model hosting. The "dev machine is a Swarm worker in the prod cluster" fact still holds structurally (the desktop is still a worker node), but nothing about Iceberg or the app database routes through it any more. - Reference implementation to copy: **`infra/ops/dbq.sh`** + `infra/ops/README.md`. ## Storage — almost certainly your warehouse Contabo Object Storage (S3-compatible), already provisioned and in use: - Endpoint `https://eu2.contabostorage.com` · bucket `era5-thermograph` · region `default` · ~2 TB · **private** (keep it that way). - ⚠️ **Contabo requires PATH-STYLE addressing.** Set `force_path_style=true` / `s3.path-style-access=true` (whatever your engine's FileIO calls it) and `region=default`. Virtual-host-style addressing **fails**. This is validated, not theoretical. - ⚠️ **The `backups/` prefix is in use** by the nightly backup jobs (prod DB + Forgejo). Put Iceberg data under its own prefix (e.g. `iceberg/` or `warehouse/`). Do not write outside your prefix. - Credentials already exist — **reuse them, don't mint new ones silently**: - SOPS vault: `infra/deploy/secrets/{prod,beta}.yaml` as `THERMOGRAPH_S3_*` (rendered to `/etc/thermograph.env` at deploy by `infra/deploy/render-secrets.sh`). - Forgejo Actions secrets: `S3_ENDPOINT`, `S3_BUCKET`, `S3_ACCESS_KEY`, `S3_SECRET_KEY`. - If read-only S3 keys are needed for the query identity, ask the operator to mint a scoped keypair rather than reusing the read-write one. ## Secrets handling - Host-side: SOPS + age. Recipient `age1xx4dzs0dxlwvkv9sjuqzsphl7lfrxannkfken374yu2qvvcte9sqzktqt2`; the private key is on each host at `/etc/thermograph/age.key` and on the operator's machine at `~/.config/sops/age/keys.txt`. - CI-side: Forgejo Actions secrets. - ⚠️ **beta gotcha:** `/etc/thermograph.env` on beta is `agent:agent 0640` — the `deploy` user **cannot read it**. If your service or job runs as `deploy` on beta, pull credentials from Actions secrets instead, or change the ownership deliberately and say so. ## Decisions you must make — and report back 1. **Catalog**: REST (Lakekeeper / Nessie / Polaris), JDBC-on-Postgres, Hive, or filesystem/hadoop — and where it runs per environment. 2. **Engine**: DuckDB + iceberg extension, Trino, Spark, or pyiceberg. Prefer the lightest that satisfies the contract; it must run containerised and headless. 3. **Warehouse location**: bucket + prefix, and whether environments are isolated by separate prefixes, namespaces, or buckets. 4. **Read-only mechanism**: scoped S3 keys? catalog RBAC? engine restricted mode? State which, and how a write is refused. 5. **Where the engine runs** for each env (local container on dev, on the box for beta/prod, or attached to the overlay) — consistent with constraint #3. ## Acceptance criteria The work is accepted when all of these are demonstrated with pasted output: - [ ] `infra/ops/iceberg.sh -c "