Guides

DataHub

Publish RawTree database and table metadata to your DataHub catalog.

DataHub

The DataHub app publishes database names, table names, column schemas, row counts, column counts, and stored table sizes from a RawTree cluster to your DataHub catalog. By default, no record values are read or copied. Optional sample publishing copies a few explicitly selected column values into DataHub. DataHub descriptions, owners, tags, and glossary terms are preserved. This integration does not enforce DataHub governance policies in RawTree.

Configure the app

  1. Select the cluster, open Apps, and install DataHub.
  2. Click the DataHub app name and enter your DataHub metadata service (GMS) endpoint, such as https://datahub-gms.example.com. This is the metadata API, not the DataHub web interface.
  3. Supply a DataHub access token (preferably for a service account) that can publish metadata. RawTree encrypts the token and never returns it in configuration responses.
  4. Optionally enter the DataHub web URL to open your catalog from RawTree.
  5. Select the databases to publish and set a stable Platform instance and Environment. These identify your catalog assets and cannot change after the configuration is saved. Use the same values as an existing rawtree-datahub source when migrating it to managed sync.
  6. Choose the sync frequency and use Test connection to check GMS connectivity. Publishing permissions are checked when metadata is actually written.
  7. Turn on Enable metadata sync and select Save and enable sync.

Installing the app alone does not start synchronization. You can save a paused configuration before the destination is reachable. Enabling sync validates the connection and queues the first run.

Follow a sync

The status page distinguishes queued work, active work, failed attempts, and the last successful sync. Sync now queues an asynchronous run; acceptance does not mean publishing has completed. Scheduled and manual runs cannot overlap for the same integration.

RawTree's existing cluster worker executes the sync through Temporal. A shared 30-second Temporal schedule discovers due integrations from durable database state and starts a separate workflow for each cluster. Configuration changes and queued runs survive API restarts and temporary Temporal outages. Short activity retries use backoff; exhausted attempts are retried after five minutes. The recurring interval starts after a complete successful run. A failed run may have updated some entries; retries publish to the same identities without deleting catalog entries. Retries rediscover the selected scope from the beginning; progress counts describe the current attempt, not a durable cursor.

Sync does not resume paused clusters. It reads schemas and table statistics from RawTree's existing table metadata and, only when enabled, the explicitly selected sample values. Row counts and stored sizes are engine-reported estimates; the sync does not run a full-table count or size scan. Statistics that the engine cannot provide are omitted; known zero values are published as zero. When a cluster is unavailable, the app waits until it is ready. A managed run is limited to five minutes, 5,000 tables, and 20,000 schema fields per table; narrow the selected databases for larger catalogs. Dynamic fields retain their native RawTree type unless type inference is enabled through the sample selection below.

Table filters

By default, all tables in selected databases are cataloged. Tables to catalog offers three modes:

  • All tables, optionally with exclusions.
  • Specific tables, an explicit database/table selection. New tables are not added automatically.
  • Name patterns, regular expressions for including current and future matching tables.

Include and exclude patterns match the complete, case-sensitive database.table name. For example, demo\.orders matches only demo.orders, demo\.orders.* matches tables starting with orders in demo, and .*\.tmp_.* matches temporary tables starting with tmp_ in any selected database. Exclusions always win. Enter one expression per line; use up to 100 patterns of at most 512 bytes each. Invalid or overly complex patterns are rejected. Backreferences and lookaround are not supported.

The catalog preview shows included and excluded tables using the same backend policy as synchronization. It refreshes as filters change and reads table names, not records. Preview supports up to 5,000 discovered tables. An empty explicit selection publishes no tables. Only tables included in metadata scope can be selected for samples; remove conflicting sample selections before saving.

Filters apply before schema and sample queries. Excluding a table stops future publishing but does not remove its existing DataHub entry or historical samples.

Optional sample values

Publish sample values is off by default. Enable it only for columns whose actual contents you intend to share with DataHub users. Select each table and column explicitly, then choose up to 5 or 10 values per column. New columns and tables are never automatically included. Schema publishing remains independent of the sampling selection.

Examples appear in DataHub's column profiles, not as complete, related rows. They are deduplicated, omit nulls, and are truncated to 256 characters. They are not a statistically representative sample and do not provide quality checks or value distributions. Each sync publishes the column count and any available row count and stored size, even when samples are off.

For a selected column whose declared type is Dynamic, RawTree also checks the runtime types in up to 1,000 rows, as the table schema view does. DataHub shows the observed concrete type or a union of observed types instead of Dynamic. If the sample has no non-null values, the type remains Dynamic. Inference is limited to selected columns and is not a guarantee about every row in the table. The type query returns only type names; it does not publish additional values. If sampling is later disabled, the next sync publishes the declared Dynamic type again.

Each value query returns at most 100 rows, and each type query at most 1,000. Both have limits of 10,000 scanned rows, 16 MiB scanned bytes, 64 MiB query memory, and five seconds of execution. Exceeding a query budget fails the run; it never falls back to an unlimited query. Select up to 100 tables and 20 columns per table. The existing five-minute overall sync limit still applies.

Sample values are stored in DataHub and governed by its access controls, not RawTree's query permissions. Disabling sampling, deselecting columns, pausing, or uninstalling stops future copies but does not delete historical samples. Remove already-published samples and profile history in DataHub if required.

Each dataset also includes a link back to its table in RawTree, where normal RawTree authentication and authorization apply. Operators should set RAWTREE_FRONTEND_URL on the existing cluster worker to the public dashboard URL.

Pause, rotate credentials, or uninstall

Use Pause to stop future publishing. To resume, enable sync in the form and save. Paste a replacement token to rotate credentials; leaving the token field blank retains the saved token. Changing the GMS endpoint requires explicitly providing its token.

Pause, configuration changes, uninstall, and cluster deletion invalidate running work before its next metadata write. An HTTPS request already in flight may finish within its 15-second timeout. Uninstall removes the configuration and encrypted token. Catalog entries remain in DataHub: this version does not remove entries when tables disappear or databases are deselected.

Validate before customer use

After enabling sync, wait for Succeeded and a new Last successful sync timestamp. Check that the selected databases and tables appear in DataHub with the expected columns, row counts, stored sizes, and RawTree links. Edit a description or tag in DataHub, run Sync now, and confirm that the edit remains after the next successful run. The app's table and column counts describe the latest attempt, not the entire catalog; inspect DataHub to confirm the customer-visible result.

Test the operating path before relying on recurring updates: pause and resume, rotate the token, and verify that a rejected token or unreachable GMS reports a failed attempt while retaining the previous successful timestamp. While the worker and Temporal are healthy, due work normally starts within one 30-second dispatch interval. An exhausted failed run is queued again after five minutes. Record the destination URL, selected scope, expected refresh interval, and the owner who will respond to failed syncs. Confirm customer network access and DataHub publishing permissions with the actual deployment; a local demo alone does not establish those conditions.

Network access

Managed publishing requires RawTree to reach GMS over HTTPS. By default, private, loopback, and link-local addresses and redirects are rejected. Customer GMS hostnames do not need to be known in advance: a publicly reachable HTTPS GMS endpoint that resolves only to public IP addresses can be configured with a DataHub access token without an allowlist entry. If DataHub is private and cannot accept connections from RawTree, run the RawTree ingestion source within that network. It reads RawTree's HTTPS metadata API and publishes locally to DataHub.

For a self-hosted deployment, the operator can explicitly permit destination origins with RAWTREE_DATAHUB_ALLOWED_ORIGINS on both the API and the existing cluster worker. For example, a local Docker demo can allow http://host.docker.internal:8080. This setting is server configuration, not something an app user can change. Tokenless publishing is permitted only for explicitly allowed origins. Saved tokens use the server's existing secret cipher and require RAWTREE_S3_STORAGE_ENCRYPTION_KEY to be configured.

API

Use a user session or OAuth credential and provide organization and cluster query parameters. Organization admins can configure and operate the app; organization members can read its redacted configuration and status.

EndpointPurpose
PUT /v1/apps/datahubInstall the app.
GET /v1/apps/datahub/configurationRead configuration and sync status.
PUT /v1/apps/datahub/configurationSave settings and desired enabled state. Include the current revision when updating.
POST /v1/apps/datahub/test-connectionCheck unsaved settings using the same configuration request shape.
POST /v1/apps/datahub/previewPreview unsaved database and table filters without saving or contacting DataHub.
POST /v1/apps/datahub/syncQueue a run. Returns 202; poll status for completion.
POST /v1/apps/datahub/pausePause without contacting the destination.
DELETE /v1/apps/datahubUninstall and remove credentials.