Extract · Transform · Orchestrate

One graph from your database to your dbt models.

GetMyData pulls from your databases and APIs, lands it where you want it, and runs your own dbt project against it — scheduled, monitored and audited. One tool, one permission model, and your extracted data never touches our servers.

9 sources · 6 destinations · 8 dbt warehouses · single sign-on included

One graph, end to end Fetch dbt
How a GetMyData pipeline fits togetherA PostgreSQL database and a REST API both feed one incremental extraction, which writes Parquet files to Amazon S3. From those files dbt builds its models and runs its tests, and the result lands in your own Snowflake warehouse.PostgreSQLsourceREST APIsourceExtractionincrementalParqueton S3dbt buildmodelsdbt testassertionsSnowflakeyour warehouse
What you get

Three things that usually need three tools.

Extraction, transformation and orchestration behind one login, one role model and one audit log — instead of stitched together across vendors.

Fetch Get the data out

Extract from a database or an API and land it in storage or your warehouse. Write SQL, reuse a saved query, or pick a table and let GetMyData build the query for you.

Parquet, CSV or JSON, optionally gzipped Incremental runs that remember where they got to Multi-step extractions that chain SQL and REST calls
dbt Run your dbt project

Point at your git repository and get scheduled jobs, run history, lineage, test results with failing rows, and docs — running your own dbt Core, not a fork of it.

GitHub, GitLab, Bitbucket, Azure DevOps or any Git remote Full-project lineage graph and a semantic layer browser Failing test rows fetched on demand, not guessed at
Pipelines Chain them together

Run an extraction, then a dbt build, on a schedule. Say what happens on success and on failure, and get told the moment something breaks.

Link a dbt source to the extraction that feeds it Per-step actions to Slack, Teams, PagerDuty, email or a webhook Seventeen alertable events, routed per person
9 Sources
PostgreSQL MySQL SQL Server Snowflake BigQuery Oracle Databricks DuckDB REST API
6 Destinations
Amazon S3 Google Cloud Storage Azure Blob Storage SFTP Windows file server Local disk
8 dbt warehouses
Snowflake BigQuery PostgreSQL Redshift Databricks DuckDB SQL Server Oracle

Every source, destination and warehouse is available today, with no per-connector charge — you are never billed for turning one on.

Built for the security review

Your data stays yours.

The questions your security team will ask, answered by how the product is built rather than by policy.

We never store your data

Extracted data is written to your storage and your warehouse. GetMyData holds the schedule, the run log and the metadata — not the rows.

A database per customer

Every workspace gets its own database with its own credentials, not a shared table with a tenant column.

No inbound firewall rules

Runners connect outbound only. When your warehouse sits inside your network, run a self-hosted runner and nothing needs to be exposed.

Single sign-on when you want it

Sign up and your first account is a local one. Once you are in, connect Entra ID, Google Workspace or any OIDC provider and your team signs in through your own directory.

Credentials encrypted at rest

AES-256-GCM secret store, or bring your own Azure Key Vault. Secrets are masked in every log, error and query preview.

Everything is audited

A queryable, exportable audit log, retained for as long as your plan allows, plus per-run logs you can stream live while a job is running.

Compute

Managed runners

We host the compute. Nothing to install, nothing to patch, and you can scale up to more replicas when a backlog builds, then let them sleep when it does not.

Metered in runner minutes — job execution time, not uptime, so an idle runner costs nothing. Prepaid, never overage: running out pauses new managed work instead of billing you more.

Compute

Self-hosted runners

Run the agent inside your own network, next to the databases it reads. It dials out to us — there is no inbound firewall rule and no VPN to maintain.

Self-hosted work consumes no runner minutes at all, so a workload you host yourself costs nothing beyond your plan.

See it against your own data.

Connect a source and run your first extraction in a few minutes — or have us walk you through it.