Reference

API

Public RawTree API reference.

API reference

This page covers the public RawTree API surface from the OpenAPI spec.

Machine-readable spec:

https://api.rawtree.com/v1/openapi.json

Base URL:

https://api.rawtree.com

Most endpoints use bearer authentication:

Authorization: Bearer rt_...

API keys are scoped to an organization and cluster, not to one database. For data endpoints, pass database=<name> in the query string or use the x-rawtree-database header. Omitting both uses the API key's stored default database. A key created without a database selector stores the logical default database.

Health

GET /health

Check service health.

Response:

{ "status": "ok" }

Databases

GET /v1/databases

List databases available in the API key's cluster. Requires a readable API key (admin, read_write, or read_only) or an authenticated organization member.

Each database includes s3_storage when it has an explicit database-level customer-owned storage override. The value is null when the database inherits the cluster's default storage. This metadata contains only bucket and path values; storage credentials are never returned.

The response also includes ClickHouse's system database for read-only inspection. RawTree manages this database and does not allow it to be deleted.

{
  "databases": [
    { "name": "default", "s3_storage": null },
    { "name": "system", "s3_storage": null },
    {
      "name": "analytics",
      "s3_storage": {
        "data": { "bucket": "acme-rawtree-data", "path": "rawtree/data" },
        "backups": { "bucket": "acme-rawtree-backups", "path": "rawtree/backups" }
      }
    }
  ]
}

POST /v1/databases

Create a database. API-key authentication requires admin permission.

{ "name": "analytics" }

To give the database its own customer-owned S3 storage, include s3_storage. Tables in the database inherit this configuration.

{
  "name": "analytics",
  "s3_storage": {
    "data": { "bucket": "acme-rawtree-data", "path": "rawtree/data" },
    "backups": { "bucket": "acme-rawtree-backups", "path": "rawtree/backups" },
    "role_arn": "arn:aws:iam::123456789012:role/RawTreeS3Access",
    "external_id": "rawtree-5b71c457-7d83-474b-9db9-6dc48a1ec042"
  }
}

DELETE /v1/databases/{database}

Delete a database and its data. Requires admin permission. Deleting a database does not revoke the cluster-wide API keys that could access it. The managed default database cannot be deleted.

Query

POST /v1/query

Execute a SQL query.

Read-only queries support SELECT, WITH, EXPLAIN, and DESCRIBE statements:

{ "sql": "SELECT * FROM events LIMIT 10" }

The Query API also supports changing a table's sorting key:

{ "sql": "ALTER TABLE events MODIFY ORDER BY (region, user.id)" }

It can also create a table, with an optional sorting key. Only the ReplicatedRawMergeTree engine is supported, and creating a table requires admin permission:

{ "sql": "CREATE TABLE IF NOT EXISTS events ENGINE = ReplicatedRawMergeTree ORDER BY (region, user.id)" }

Other engines, column definitions, engine arguments, AS SELECT, and clauses other than ORDER BY are rejected with a 400. The error hint includes a supported statement you can copy. Creating a table that already exists returns 409 unless the statement uses IF NOT EXISTS.

To materialize query results into another table, such as a rollup or a cleaned copy of raw events, use INSERT INTO ... SELECT:

{
  "sql": "INSERT INTO daily_clicks SELECT toDate(timestamp) AS day, count() AS clicks FROM events GROUP BY day"
}

INSERT INTO ... SELECT has these limits:

  • The target table must already exist in the selected database. Create it with CREATE TABLE, POST /v1/tables, or by inserting data first.
  • The API key needs both read and write permission.
  • URL table functions (url() and urlCluster()) are rejected.
  • SETTINGS clauses are passed to ClickHouse for validation and execution.
  • INSERT ... VALUES and INSERT ... FORMAT are not supported; send JSON rows to POST /v1/tables/{table} instead.
  • RawTree supplies a default max_execution_time of 30 seconds. An explicit SETTINGS clause can override it.

The response is an empty result set. statistics.written_rows reports how many rows were inserted:

{
  "meta": [],
  "data": [],
  "rows": 0,
  "statistics": { "elapsed": 0.42, "rows_read": 120000, "bytes_read": 9600000, "written_rows": 31 }
}

The optional format field selects the ClickHouse response format. It defaults to JSON, and supports JSONEachRow, JSONEachRowWithProgress, JSONCompact, CSV, CSVWithNames, TSV, TSVWithNames, and TabSeparated. For formats other than JSON, the response body is returned in the requested raw format. For example:

{
  "sql": "EXPLAIN SELECT 1",
  "format": "JSONEachRowWithProgress"
}

JSONEachRowWithProgress returns newline-delimited JSON records, including progress, metadata, and row records.

Default JSON response shape:

{
  "meta": [{ "name": "action", "type": "String" }],
  "data": [{ "action": "click" }],
  "rows": 1,
  "statistics": { "elapsed": 0.001, "rows_read": 1, "bytes_read": 128 },
  "hints": []
}

Row policies

Organization administrators and admin API keys can create row policies through /v1/query. A policy filters the rows that an existing database role can read from one table. Select the table's database with database=<name> and send each statement in a separate request:

CREATE ROW POLICY acme_events ON events
FOR SELECT USING __raw_data.tenant_id::String = 'acme'
TO analytics_reader;

The role must already exist and have a SELECT grant on the table. Create an API key with database_roles: ["analytics_reader"] to query as that role; see database-role keys. Permission-based keys and user sessions do not assume this role.

Supported statements are CREATE ROW POLICY [IF NOT EXISTS | OR REPLACE] and DROP ROW POLICY [IF EXISTS]. OR REPLACE replaces the complete definition, including its role targets:

CREATE ROW POLICY OR REPLACE acme_events ON events
FOR SELECT USING __raw_data.tenant_id::String = 'acme' AND __raw_data.environment::String = 'production'
TO analytics_reader;
DROP ROW POLICY IF EXISTS acme_events ON events;

An optional ON CLUSTER 'default' can appear immediately before or after ON database.table. With ON CLUSTER, the database must be explicit and match the selected database; unqualified targets are rejected because distributed access DDL can otherwise resolve a replica's default database. Omit ON CLUSTER when using replicated access storage; otherwise include it to apply the statement to every current replica. The API preserves this choice and reports distributed DDL errors. With replicated access storage, plain CREATE or DROP combined with ON CLUSTER can race between replicas; use IF NOT EXISTS or IF EXISTS when sending those forms.

Policies default to AS PERMISSIVE; AS RESTRICTIVE is also supported before or after the USING filter. Matching permissive policies combine with OR, then matching restrictive policies combine with AND. A restrictive policy needs a matching permissive policy to allow any rows.

RawTree currently allows readers without any matching policy to keep their existing access. Creating a policy for one role does not restrict other roles, permission-based keys, or user sessions. This depends on the server setting access_control_improvements.users_without_row_policies_can_read_rows; disabling it denies rows to readers without a matching permissive policy. Plan policies for every role that needs restricted access. Policies filter SELECT results; they do not constrain inserts or provide isolation from users who can modify tables. Use read-only roles for row-restricted readers.

Each request accepts one policy on one table in the selected user database, a scalar filter expression, and 1–32 distinct existing roles. Cross-database targets, wildcards, subqueries, user targets, ALL, ALL EXCEPT, NONE, CURRENT_USER, ALTER ROW POLICY, and access-storage selectors are unsupported. Role names that collide with database user names are rejected. System databases and the internal organization do not support policy management.

Workload scheduling

Separate interactive queries and batch processing into named workloads with CPU and disk I/O scheduling policies. Organization administrators and admin API keys can manage definitions through /v1/query:

CREATE WORKLOAD all ON CLUSTER 'default' SETTINGS max_concurrent_threads = 8;
CREATE WORKLOAD default ON CLUSTER 'default' IN all;
CREATE WORKLOAD interactive ON CLUSTER 'default' IN all SETTINGS weight = 3;
CREATE WORKLOAD batch ON CLUSTER 'default' IN all SETTINGS weight = 1;
CREATE RESOURCE cpu ON CLUSTER 'default' (MASTER THREAD, WORKER THREAD);

Send each statement in its own request, using the sql field. Definitions are independent of the selected database. Choose replication according to the cluster's workload storage:

  • Local workload storage: include ON CLUSTER 'default' after the object name, as in the example, to apply the definition to the cluster's current replicas. Without it, the statement affects only the server that receives the request.
  • Keeper-backed workload storage (workload_zookeeper_path): omit ON CLUSTER; the engine replicates definitions automatically and rejects ON CLUSTER.

The API preserves this choice. Other cluster names, multiple statements, and backslash escapes in identifiers are rejected. The internal organization does not support scheduling management.

Supported management statements:

  • CREATE WORKLOAD [IF NOT EXISTS] <name> [IN <parent>] [SETTINGS ...]
  • CREATE OR REPLACE WORKLOAD <name> [IN <parent>] [SETTINGS ...]
  • DROP WORKLOAD [IF EXISTS] <name>
  • CREATE RESOURCE [IF NOT EXISTS] <name> (<operations>)
  • CREATE OR REPLACE RESOURCE <name> (<operations>)
  • DROP RESOURCE [IF EXISTS] <name>

Resource operations are MASTER THREAD, WORKER THREAD, READ DISK <name>, WRITE DISK <name>, READ ANY DISK, and WRITE ANY DISK. Memory reservation and query-slot resources are not supported by this API. Workload settings accept numeric or string literals and an optional FOR <resource> per setting; the engine validates setting names, values, and hierarchy constraints. For example:

{
  "sql": "CREATE OR REPLACE WORKLOAD batch ON CLUSTER 'default' IN all SETTINGS weight = 1, max_concurrent_threads = 2"
}

Replacement supplies the full definition. Remove children before dropping their parent, and remove resource-specific workload settings before dropping the referenced resource. Successful management returns the same empty result shape as other DDL. A replica error or distributed DDL timeout returns an error; cluster DDL is not transactional, so inspect the definitions before retrying. With local workload storage, provision the same definitions when adding replicas. Keeper-backed storage supplies definitions to new replicas automatically; configure it in the cluster infrastructure and omit ON CLUSTER from management requests.

Select a leaf workload for a read query or INSERT INTO ... SELECT using a SQL SETTINGS clause:

{
  "sql": "SELECT count() FROM events SETTINGS workload = 'interactive'"
}

Read-only and custom database-role keys can select a workload while retaining their existing data permissions. Workload selection does not grant access to tables.

Scheduling only takes effect when resources and the referenced workload exist. The engine's default workload is default; with throw_on_unknown_workload disabled, unknown names bypass scheduling. This endpoint does not change that server policy. For enforced CPU scheduling, operators must also constrain use_concurrency_control so callers cannot disable it. CPU limits apply per server, not as one cluster-wide budget. See ClickHouse workload scheduling for resource configuration and enforcement.

Inspect definitions and scheduling activity with SELECT queries against system.workloads, system.resources, and system.scheduler using a key with access to those system tables.

Workload settings profiles

Assign a default workload to a database role so its API keys do not need a SETTINGS workload clause on every query. Organization administrators and admin API keys can create, replace, and drop workload settings profiles through /v1/query. Create the role and workload first, then send this SQL in the sql field:

CREATE SETTINGS PROFILE interactive_profile
SETTINGS workload = 'interactive'
TO app_reader;

API keys with database_roles: ["app_reader"] inherit the default. Without a constraint, callers can choose another workload in SQL. To lock the workload and prevent callers from disabling CPU concurrency control, replace the profile:

CREATE SETTINGS PROFILE OR REPLACE interactive_profile
SETTINGS workload = 'interactive' READONLY, use_concurrency_control = 1 CONST
TO app_reader;

READONLY and CONST lock the individual setting; they do not grant or remove table permissions. Profiles apply to existing and newly created API keys using the selected roles. Replacement supplies the entire definition, including the role list. Custom-role keys retain their existing Query API restrictions.

Supported statements are CREATE SETTINGS PROFILE [IF NOT EXISTS], CREATE SETTINGS PROFILE OR REPLACE, and DROP SETTINGS PROFILE [IF EXISTS]. Submit one profile and one statement per request. Profile settings are limited to workload (a nonempty string) and use_concurrency_control (0 or 1), with optional READONLY or CONST constraints. Creation requires TO followed by 1–32 distinct existing roles without same-named database users. User targets, ALL, NONE, inheritance, other settings, and ALTER SETTINGS PROFILE are not supported. Profile management is not available to non-admin callers, workflows, or the internal organization.

An optional ON CLUSTER 'default' after the profile name retains ClickHouse's native access-storage behavior. Settings profiles use access storage, which is separate from workload storage: with replicated access storage, profiles replicate automatically; local access storage needs ON CLUSTER for all current replicas. Omit ON CLUSTER with replicated access storage to avoid duplicate creation or deletion as access entities replicate between nodes. The API does not add the clause. Other cluster names are rejected.

DROP SETTINGS PROFILE IF EXISTS interactive_profile;

See ClickHouse settings profiles for the underlying SQL syntax. Workloads and resources must still be configured for scheduling to take effect.

Logs

GET /v1/logs

List platform API requests captured by RawTree's internal request spans. Query execution history belongs to the Queries screen; this endpoint focuses on HTTP request activity.

Query parameters:

ParameterDescription
start_timeInclusive lower bound.
end_timeInclusive upper bound.
limitRows to return. Default 50, max 200.
offsetRows to skip.
searchFree-text search across log ID, method, path, status, source, user agent, and errors.
methodsComma-separated HTTP methods, such as GET,POST.
status_codesComma-separated HTTP status codes, such as 200,404,500.
sourcesComma-separated request sources: ui, cli, or api.
user_agentExact user-agent value.

Each log includes its span-based log ID, creation time, method, URL path, matched API route, status code, source, duration, user agent, client version, request size, trace ID, database, and structured error details when the request failed. A bounded textual request body may be included for supported requests (up to 8 KiB); response body content is not recorded.

Tables

GET /v1/tables

List tables in the selected database.

POST /v1/tables

Create a table. Requires admin permission.

{ "name": "events", "sorting_key": "region, user.id" }

sorting_key is optional. Omit it, or send an empty string, and the table picks a sorting key per part from the ingested data. The string contains SQL expressions in key order, separated by commas. A bare name such as user.id is read as a path into the ingested JSON.

GET /v1/tables/{table}

Describe a table. The response includes sorting_key, which is empty when the table picks a key per part.

PATCH /v1/tables/{table}

Update a table's ClickHouse configuration. Requires admin permission. Only the fields you send are changed; omitted fields are left as they are.

{ "sorting_key": "timestamp" }

Functions and multiple expressions use the same string field. Commas inside function arguments remain part of that function:

{ "sorting_key": "ifNull(cityHash64(host, instanceId), 0)" }

Bare column names inside functions resolve to fields in the ingested JSON. Use SQL identifier quotes for unusual column names, such as `user-agent`. Create, update, and describe responses return sorting_key as a string; an empty string indicates automatic sorting.

The new sorting key applies to newly inserted parts and wins later merges, so existing parts are re-sorted in the background rather than rewritten by the request. The key must contain at least one column or expression; a sorting key cannot be removed once set. RawTree checks expression syntax before sending the key to the engine. An unusable key, such as one that repeats a column, is answered with the engine's own error.

POST /v1/tables/{table}

Insert data. Send one JSON object or an array of JSON objects.

[{ "action": "click", "user": "alice" }]

Optional transform for JSON body inserts:

POST /v1/tables/traces?transform=otlp-traces

Transforms flatten known source formats before insert.

See Transforms for supported input shapes and emitted rows.

Supported transforms:

  • otlp-traces
  • otlp-logs
  • otlp-metrics
  • cloudwatch-logs
  • cloudtrail
  • firehose

For AWS Firehose HTTP endpoint delivery, use:

POST /v1/tables/events?transform=firehose
X-Amz-Firehose-Access-Key: <rawtree-api-key>

Firehose records must contain base64-encoded data in records[].data. If the decoded value is a JSON object, RawTree inserts it as one row. If it is an array of JSON objects, RawTree inserts one row per object. If it is UTF-8 TSV text and columns is present, RawTree inserts one row per TSV line using those column names. If it is UTF-8 TSV text without columns, RawTree inserts one row with the decoded text in data. Other decoded formats return 400. Successful Firehose inserts return {requestId,timestamp}.

URL ingest uses query parameters:

POST /v1/tables/events?url=https%3A%2F%2Fexample.com%2Fevents.jsonl

The request stays open until the native ClickHouse URL import finishes. Success returns 200 with JSON such as {"inserted":1000}; inserted is null if ClickHouse does not provide a row count. Errors return a normal HTTP error response. There is no progress event stream.

Optionally pass query_id with a URL import to correlate it with ClickHouse query logs. If omitted, ClickHouse generates the ID. In both cases, the ID is returned in the X-ClickHouse-Query-Id response header, matching the query API. It is not added to the JSON body.

Transforms are not supported with URL inserts. If you use ?url=, transform the data before hosting it.

DELETE /v1/tables/{table}

Delete a table. Requires admin permission.

API keys

GET /v1/keys

List API keys for the current cluster. Requires an admin API key or JWT.

POST /v1/keys

Create a cluster-wide API key. Requires an admin API key or JWT. The database selected on this request is stored as the key's default for later requests that omit a database selector; if none is selected, RawTree stores default.

{ "name": "my-agent", "permission": "read_write" }

Valid permissions:

  • admin
  • read_write
  • write_only
  • read_only

DELETE /v1/keys/{id_or_token}

Delete an API key by UUID or full rt_... token.

Clusters

Cluster endpoints require an authenticated user session or OAuth access token. Pass the organization name as organization on every endpoint except the size catalog. Organization members can read clusters and metrics; organization admins are required to create, update, pause, resume, verify storage, or delete clusters.

Cluster responses use this shape:

{
  "id": "9f30c31c-74a7-4a08-aab0-494689ab5b31",
  "name": "production",
  "created_at": "2026-08-26 10:00:00+00",
  "status": {
    "phase": "ready",
    "ready": true,
    "message": null
  },
  "resources": {
    "shards": 1,
    "replicas": 2,
    "cpu_cores_per_replica": 2.0,
    "memory_bytes_per_replica": 8589934592,
    "minimum_size": { "cpu_cores": 2.0, "memory_bytes": 8589934592 },
    "maximum_size": { "cpu_cores": 64.0, "memory_bytes": 274877906944 }
  },
  "can_pause": true,
  "can_resume": false,
  "s3_storage": null,
  "database_s3_access": null,
  "idle_timeout_minutes": 15
}

resources can be null while resource information is unavailable. Use status.ready, can_pause, and can_resume instead of inferring lifecycle capabilities from the phase name.

GET /v1/clusters

List the organization's permanent clusters and their current status.

GET /v1/clusters?organization=acme

The response contains an array of cluster objects:

{
  "clusters": []
}

GET /v1/clusters/{cluster_id}

Get one cluster by the UUID returned from the cluster list.

GET /v1/clusters/9f30c31c-74a7-4a08-aab0-494689ab5b31?organization=acme

GET /v1/clusters/sizes

List the supported per-replica resource sizes, replica limits, and default autoscaling bounds. Call this endpoint before creating a cluster and pass one of its cpu_cores and memory_gib pairs back to the create endpoint.

{
  "min_number_of_replicas": 1,
  "max_number_of_replicas": 2,
  "sizes": [
    { "size": "large", "cpu_cores": 2, "memory_gib": 8 },
    { "size": "xlarge", "cpu_cores": 4, "memory_gib": 16 }
  ],
  "default_min_size": { "size": "large", "cpu_cores": 2, "memory_gib": 8 },
  "default_max_size": { "size": "xlarge", "cpu_cores": 4, "memory_gib": 16 }
}

The size label is for display. The resource pair is the stable create input.

POST /v1/clusters

Create a cluster. name, replicas, and size are required. autoscaling, idle_timeout_minutes, s3_storage, and database_s3_access are optional.

POST /v1/clusters?organization=acme
{
  "name": "production",
  "replicas": 2,
  "size": { "cpu_cores": 2, "memory_gib": 8 },
  "autoscaling": {
    "min_size": { "cpu_cores": 2, "memory_gib": 8 },
    "max_size": { "cpu_cores": 8, "memory_gib": 32 }
  },
  "idle_timeout_minutes": 15
}

Set idle_timeout_minutes to 0 to disable automatic idling, or use a value from 15 through 43,200 minutes. Omit it to use the configured default. When autoscaling is omitted, the size catalog's default_min_size and default_max_size are used. The initial size must fall within those bounds.

To use customer-owned S3 storage, include s3_storage:

{
  "name": "production",
  "replicas": 2,
  "size": { "cpu_cores": 2, "memory_gib": 8 },
  "s3_storage": {
    "data": { "bucket": "acme-rawtree-data", "path": "rawtree/data" },
    "backups": { "bucket": "acme-rawtree-backups", "path": "rawtree/backups" },
    "role_arn": "arn:aws:iam::123456789012:role/RawTreeS3Access",
    "external_id": "rawtree-5b71c457-7d83-474b-9db9-6dc48a1ec042"
  }
}

Tables created normally use the cluster's configured storage. If cluster creation returns an ambiguous response, list clusters and look for the requested name before retrying.

To keep RawTree-managed cluster defaults while allowing databases to use their own customer-owned buckets, include database_s3_access when the cluster is created:

{
  "name": "production",
  "replicas": 2,
  "size": { "cpu_cores": 2, "memory_gib": 8 },
  "database_s3_access": {
    "external_id": "rawtree-5b71c457-7d83-474b-9db9-6dc48a1ec042",
    "database_bucket_tag": "rawtree-dc1fb78e-341d-48f8-a183-9ecea12eb41d"
  }
}

This capability is immutable after cluster creation. Tag every database bucket and IAM role that the cluster may use with rawtree.com/cluster=<database_bucket_tag>. When s3_storage and database_s3_access are both present, they must use the same external_id.

POST /v1/clusters/verify-s3-access

Verify that RawTree can assume the configured role and access both S3 destinations before creating the cluster. The s3_storage object has the same shape as the create request.

POST /v1/clusters/verify-s3-access?organization=acme
{
  "s3_storage": {
    "data": { "bucket": "acme-rawtree-data", "path": "rawtree/data" },
    "backups": { "bucket": "acme-rawtree-backups", "path": "rawtree/backups" },
    "role_arn": "arn:aws:iam::123456789012:role/RawTreeS3Access",
    "external_id": "rawtree-5b71c457-7d83-474b-9db9-6dc48a1ec042"
  }
}
{ "verified": true, "message": "S3 access verified." }

PATCH /v1/clusters/{cluster_id}

Update the cluster name, automatic idle timeout, or both. Omitted properties are unchanged.

PATCH /v1/clusters/9f30c31c-74a7-4a08-aab0-494689ab5b31?organization=acme
{
  "name": "primary",
  "idle_timeout_minutes": 60
}

Use idle_timeout_minutes: 0 to disable automatic idling.

POST /v1/clusters/{cluster_id}/stop

Request that a running cluster pause. The response is the updated cluster object; use its status and can_pause/can_resume fields to follow the transition.

POST /v1/clusters/{cluster_id}/resume

Request that a paused cluster resume. The response is the updated cluster object. A conflicting lifecycle transition returns 409.

DELETE /v1/clusters/{cluster_id}

Permanently delete a cluster.

{ "deleted": true }

GET /v1/metrics

Query internal cluster and connector telemetry using a first-party session or OAuth access token. API keys are not accepted. Select a scope using either organization_id and optional comma-separated cluster_id UUIDs, or organization and cluster names.

GET /v1/metrics?organization=acme&cluster=production&metric=rawtree.table.storage.size&start_time=2026-09-22T11%3A00%3A00Z&end_time=2026-09-22T12%3A00%3A00Z

metric accepts comma-separated metric names. Existing internal metrics include rawtree.cluster.cpu.allocated, rawtree.cluster.memory.allocated, rawtree.table.storage.size, rawtree.table.rows, messaging.client.consumed.messages, kafka.consumer_group.lag, and the rawtree.connector.* event, error, buffer, availability and latency metrics. Allocated capacity is distinct from observed CPU/memory usage.

filters accepts a URL-encoded JSON object of exact attribute matches, such as {"rawtree.connector.id":"<id>","rawtree.connector.revision":"2"} or {"db.namespace":"analytics","db.collection.name":"events"}. Values must be strings. Scope parameters enforce access; attribute filters cannot change tenant ownership. Historical connector revisions are available when explicitly selected.

For live or trailing-range queries, pass window_seconds (1 through 604800) instead of start_time and end_time. The server resolves both boundaries using its UTC clock and returns the resolved start_time and end_time in the response. Use that returned end time for freshness checks rather than the browser clock. Relative and explicit ranges cannot be combined.

For stored metrics, start_time is inclusive and end_time exclusive, both RFC 3339 timestamps, with a maximum seven-day interval. include_baseline=true also requests a recent preceding observation for rate calculations. Baselines are optional and may be absent. Long or dense ranges sample complete observations; returned points are observations, not bucket sums, averages or rates.

The response contains resources and partial_errors. Each resource identifies its organization, cluster and producer, with an instrumentation scope and a metrics array. Stored series have source: "internal.metrics", name, unit, type, and points. Points retain attributes and OTLP fields such as timeUnixNano, asInt/asDouble, or histogram count, sum, explicitBounds and bucketCounts. The 64-bit measurement integers and nanosecond timestamps use decimal strings. This is a RawTree query response organized around the OTel data model, rather than an OTLP export envelope. Known connector sums/histograms report cumulative temporality; sums also report monotonicity. Calculate counter deltas per series with reset handling. Missing observations are not zero measurements.

The existing max_cpu_cores and max_memory_bytes chart series remain supported and still query engine system tables. They are the default when metric is omitted. Attribute filters and baselines apply to stored metric requests only.

The connector-specific /v1/connectors/{connector_id}/metrics route has been removed. Read connector state from /v1/connectors/{connector_id} and query its metrics here. Cluster and workflow metric routes remain available while their remaining data sources are migrated.

GET /v1/clusters/{cluster_id}/metrics

Return time-series metrics for a cluster UUID. organization, start_time, and end_time are required query parameters. metrics is an optional comma-separated list; omitting it returns every supported metric.

GET /v1/clusters/9f30c31c-74a7-4a08-aab0-494689ab5b31/metrics?organization=acme&start_time=2026-08-26T09%3A00%3A00Z&end_time=2026-08-26T10%3A00%3A00Z&metrics=total_read_queries,max_cpu_cores

Supported metric names:

  • total_read_queries
  • total_write_queries
  • max_cpu_cores
  • max_memory_bytes

The response returns one series per requested metric. Every series describes its unit, aggregation, scope, source, and data points. A point can include a replica name when the metric is replica-scoped.

{
  "metrics": [
    {
      "name": "max_cpu_cores",
      "label": "Max CPU",
      "description": "Maximum ClickHouse process CPU cores observed in each bucket.",
      "unit": "cores",
      "aggregation": "max",
      "scope": "cluster",
      "source": "system.metric_log",
      "points": [
        { "time": "2026-08-26T09:00:00Z", "value": 1.5, "replica": "rawtree-0" }
      ]
    }
  ]
}

GET /v1/clusters/metrics

Return the same metrics response using a cluster name instead of its UUID. Pass the required organization and cluster query parameters together with start_time, end_time, and the optional metrics list.

Workflows

Create and manage scheduled SQL workflows through /v1/workflows. The Workflows API reference documents creation, updates, pause/resume, deletion, execution logs, and metrics. See the workflow guide for SQL behavior and sinks.

Apps

Apps are installed per cluster, and clusters start with no apps installed. App management requires an authenticated user session or OAuth access token; cluster API keys cannot manage apps.

GET /v1/apps

List the available apps and their installation state. Any member of the organization can list apps. Pass the organization and cluster names in the query string:

GET /v1/apps?organization=acme&cluster=production
{
  "apps": [
    { "id": "opentelemetry", "name": "OpenTelemetry", "installed": true },
    { "id": "prometheus", "name": "Prometheus", "installed": false }
  ]
}

This endpoint reads platform metadata only, so it remains available while the cluster is provisioning, paused, stopped, or otherwise unavailable.

PUT /v1/apps/{app_id}

Install an app. Organization admin access is required. The request has no body and is idempotent:

PUT /v1/apps/prometheus?organization=acme&cluster=production
{ "id": "prometheus", "name": "Prometheus", "installed": true }

Installation enables the app's native endpoints; it does not change the cluster lifecycle state or guarantee that the cluster is ready.

DELETE /v1/apps/{app_id}

Uninstall an app. Organization admin access is required. The request has no body and is idempotent:

DELETE /v1/apps/prometheus?organization=acme&cluster=production
{ "id": "prometheus", "name": "Prometheus", "installed": false }

Uninstallation disables the app's native endpoints; it does not change the cluster lifecycle state. The Apps catalog remains available while the cluster is stopped.

The following sections document the APIs and protocols implemented by each app.

OpenTelemetry app

RawTree accepts native OpenTelemetry Protocol ingest over OTLP/HTTP and OTLP/gRPC. For setup instructions, SDK environment variables, Collector config, and a smoke test, see the OpenTelemetry guide.

The OpenTelemetry app must be installed on the target cluster before these native endpoints can be used. The generic table insert transforms, such as POST /v1/tables/traces?transform=otlp-traces, do not require the app.

Both paths apply the built-in OpenTelemetry transforms and write to the default signal tables.

Native OTLP endpoints write to the default traces, logs, and metrics tables unless you provide a signal-specific destination header: x-rawtree-traces-table, x-rawtree-logs-table, or x-rawtree-metrics-table. For OTLP/HTTP, select the database with the database query parameter. For OTLP/gRPC, use the x-rawtree-database metadata header. If the request omits a selector, RawTree uses the API key's stored default database.

POST /otlp/v1/traces

Ingest OTLP/HTTP traces into the traces table, or into the table named by x-rawtree-traces-table.

POST /otlp/v1/logs

Ingest OTLP/HTTP logs into the logs table, or into the table named by x-rawtree-logs-table.

POST /otlp/v1/metrics

Ingest OTLP/HTTP metrics into the metrics table, or into the table named by x-rawtree-metrics-table.

These endpoints accept application/json and application/x-protobuf OTLP export payloads. Successful requests normally return an empty OTLP export response: {} for JSON requests or an empty protobuf message for protobuf requests. If RawTree accepts the export but drops invalid signal records, the response uses the OTLP partialSuccess shape with rejectedSpans, rejectedLogRecords, or rejectedDataPoints plus an errorMessage.

OTLP/HTTP request bodies can be gzip-compressed with Content-Encoding: gzip. Request bodies are limited to 100 MiB after decompression; larger exports return 413.

OTLP/gRPC

Send OTLP/gRPC export requests to the standard collector services:

SignalService
Tracesopentelemetry.proto.collector.trace.v1.TraceService/Export
Logsopentelemetry.proto.collector.logs.v1.LogsService/Export
Metricsopentelemetry.proto.collector.metrics.v1.MetricsService/Export

Use https://api.rawtree.com as the OTLP endpoint, or http://localhost:4317 with the local Docker Compose stack.

Use bearer authentication with a cluster API key:

Authorization: Bearer rt_...

Prometheus app

RawTree exposes Prometheus remote write and query-compatible endpoints under /prometheus/api/v1. The Prometheus app must be installed on the target cluster before any endpoint in this group can be used. See the Prometheus guide for configuration and supported endpoints.

The Splunk app enables both the Splunk HEC-compatible ingestion endpoints and RawTree's platform-native SPL search API for a cluster. SPL searches run inside RawTree against the selected database's splunk_events table; they do not install or require a Splunk search provider, command, or add-on.

GET /v1/spl/capabilities

Return the static SPL language manifest for this backend version. The response describes supported commands and their valid pipeline positions, functions, operators, keywords, literals, timechart span units, built-in fields, and compiler limits. Clients can fetch it once and compute contextual completion locally; no organization, cluster, database, or authentication selector is required. Database column names are intentionally separate and remain available through the table-description API.

POST /v1/spl/search

Execute a supported, read-only SPL search. A readable cluster API key uses its bound organization, cluster, and database when those selectors are omitted. With a user or OAuth token, pass the organization, cluster, and database. As on other data endpoints, x-rawtree-database can select the database instead of the database query parameter.

POST /v1/spl/search?organization=acme&cluster=production&database=logs
Authorization: Bearer rt_...
{
  "spl": "index=bot_traffic endpoint=\"/api/login\" action=blocked | sort - _time",
  "earliest": "-30m",
  "latest": "now",
  "include_field_statistics": true,
  "query_id": "optional-client-query-id"
}

Time bounds accept non-negative epoch seconds, now, or a negative relative value such as -30m, -24h, or -7d. RawTree applies the same read-only query safeguards and cancellation ownership rules as the SQL query API. Cancel a running request with POST /v1/query/cancel using the same query_id, organization, and database scope.

The supported SPL subset covers base searches and search/where boolean filters; positive fields/table; rename; bounded eval; stats with count, dc, sum, avg, values, min, max, earliest, and latest plus optional BY; chart; sort; head; multi-field dedup; and timechart with multiple aggregates and an optional split field. count(eval(<condition>)) is supported; other aggregate eval expressions are not. Unsupported or ambiguous commands return 400 instead of being passed through as SQL.

stats ... BY a b produces rows grouped by both fields. chart ... OVER a BY b (also chart ... BY a b) produces rows for a and separate series columns for b. Split timechart produces the same wide series layout with _time as the first column. With multiple aggregates, series names are <aggregate>:<value>; an explicit AS alias replaces the aggregate name. Quote output field names containing punctuation, for example table _time 'mean:US'. These columns exist before downstream SPL stages run. The response marks this layout with result_format: "native_columns"; older snapshots without this marker retain their legacy rendering behavior.

Both chart commands support limit=N (default 10), limit=topN, limit=bottomN, limit=0 (all series within the safety bound), useother=true|false, and usenull=true|false. Excluded series are combined as OTHER and null split values appear as NULL by default. Series ranking uses the sum of per-row aggregate values for one aggregate, or event frequency for multiple aggregates. The compiler discovers series through bounded read-only SQL, then aggregates the selected series through the same scoped query execution and cancellation path. At most 20 aggregates and 100 output series columns are allowed; excessive series or colliding output names return 400 rather than silently dropping data.

timechart accepts span intervals in s, m, h, d, or w, from 1 second through 31 days, and bins=N (default 100) for automatic span selection. Automatic spans use a bounded interval ladder across the requested window, not every Splunk calendar/span option. cont=true fills missing buckets by default; cont=false returns only observed buckets. Counts fill with zero; absent sums and averages remain null. More than 10,000 continuous buckets require a larger span or smaller window. minspan, calendar-month spans, custom NULL/OTHER labels, and explicit series where clauses are not supported.

The JSON response includes the standard query meta, data, rows, and statistics fields, plus the original spl, a bounded, chronologically ordered timeline, a result_kind of events or statistics, a truncated flag, and an entry for every pipeline segment in source order. Execution uses Smart mode: event-preserving searches produce events and a timeline; transforming searches (stats, chart, timechart, and table) produce statistical results with an empty timeline and no raw_events or event field_statistics. There is no Fast/Verbose mode request option. Timeline buckets for event searches are zero-filled across the requested window (or the recent default window when no time bounds are supplied), with granularity selected to keep the series bounded. Each stage has command, expression, output_rows, and truncated fields. output_rows is the row count after that stage. It is capped at 10,000; truncated: true means a 10,001st row proved that the real count is larger, while exactly 10,000 rows remain exact with truncated: false.

Set include_field_statistics to true to include field_statistics for the event fields produced by the search. Each entry contains the inferred field type, total matching rows, non-null value count, coverage, exact distinct-value count, the five most common values with exact counts, and the source indexes or tables in which the field was found. These statistics are calculated in ClickHouse across the complete filtered event population; the 10,000-row result limit applies only to rows returned in data. This option is ignored for transforming searches in Smart mode. RawTree's SPL editor uses these common values as field-specific suggestions on the right-hand side of search and where comparisons. They are ranked hints, not an exhaustive list of valid values.

RawTree measures cardinality-changing stages and carries exact counts across row-preserving stages. Measurement is limited to the first 12 stages and to the generated-query size budget, so output_rows is null when a stage cannot be measured safely. Clients should calculate percentages only for consecutive stages whose counts are non-null and untruncated, where the previous count is positive and the current count does not exceed it. In that case, kept percent is current / previous * 100, affected rows is previous - current, and affected percent is 100 - kept percent; percentage values use the 0–100 range.

Search execution is synchronous: the response is returned after the result, optional raw-event and field-statistics queries, and the bounded analysis query for timeline and stage feedback complete. Every query uses the same authorization, database scope, read-only validation, execution limits, and cancellation ownership.

Saved searches

Saved searches are scoped to an organization and cluster. Private searches are visible only to their creator; shared searches are available to authenticated members of the same organization while the Splunk app is installed:

EndpointPurpose
GET /v1/apps/splunk/searchesList and filter saved searches.
POST /v1/apps/splunk/searchesCreate a saved search.
GET /v1/apps/splunk/searches/{search_id}Get one saved search.
PATCH /v1/apps/splunk/searches/{search_id}Update a saved search.
DELETE /v1/apps/splunk/searches/{search_id}Delete a saved search.

Pass organization and cluster on every saved-search request. List requests also accept search, created_by, database, and visibility filters. New searches default to private; set visibility to shared to share one with other authorized organization members. The complete record persists its SPL, time window, selected table or chart visualization, time-series span in minutes, and whether the result table remains available. The list returns metadata for at most 200 searches and sets truncated when additional matches exist; fetch one search to read its complete configuration. Uninstalling the Splunk app disables HEC, SPL execution, and saved-search access but preserves saved-search records and ingested data. Reinstalling the app makes them available again; SQL access is unaffected.