connector.taps[]. You do not set them.
schema (and outputSchema on your overlay, if you
narrowed it), plus dataKind, identityDirectory, and
responseExtraction.primaryKey.
From flags to a queryable table
responseExtraction.primaryKey is how ingestion recognizes a fetched
record (almost always id). Lake uniqueness is the x-transformDedupKey
set. They usually name the same field. Child tables often use a composite
(for example type + value on a nested Google user email).
If no dedup key is declared, the lake falls back to a column named id.
Names in SQL
Address:providers.{connector}.{table}. {connector} is the connector
slug. {table} is the tap name, lowercased. Nested object arrays become
extra tables named {tap}__{path} (GitHub issue labels are
providers.github.issues__labels). Confirm names in the
catalog before you query; not every
nested object explodes.
Column names are the schema property names, ASCII-lowercased. Nested
fields that flatten use __ between segments:
SELECT updatedAt against Linear issues will not find the column. Use
updatedat.
Dedup keys
Every current row in a provider table is unique on its dedup key columns. Linearissues and GitHub pull_requests both key on id.
You can join or filter on that column and expect one row per issue or PR.
When a key is composite, you need every part. A nested Google emails
child table keys on address; a nested external-id child keys on
value and type. Look at x-transformDedupKey on that nested
type, not only on the root.
Do not SELECT DISTINCT id to “clean up” a provider table. Dedup
already happened. If you see two rows with the same id, you are
probably looking at a child table (parent id repeated per nested
element) or you joined without the full key.
Ordering keys
A tap may declare at most onex-transformOrdering field. When
ingestion writes a newer payload for the same dedup key, the row with
the greater ordering value is the one you query.
Linear issues: updatedAt (SQL updatedat). GitHub pull requests:
updated_at. Directory taps such as Linear users also order on
updatedAt so a renamed user replaces the previous row.
dataKind is how those versions are treated:
deletionSemantics tells you whether a missing source record disappears
from the table. undetectable (Linear issues, GitHub PRs) means a
deleted item can remain until something else updates that key.
by_absence on a snapshot tap means a row gone from the latest extract
is gone from the table. tombstone means the source sends an explicit
delete marker.
Identity keys
Identity flags mark which columns are about a person or account, not which row is unique. Linearusers is the full set on one tap:
id is both the row’s dedup key and the Linear account id. email is
how you join that person to Google primaryemail or workspace.users.email.
name is for matching, not for joins.
Taps with identityDirectory: true (Linear users, GitHub members,
Google users) contribute to providers.identity.*. That family appears
only after at least one directory tap has synced.
Association uniqueness is
(connector_id, connector_tap_id, account_id).
Join Google’s id (account id) through associations; do not join two
providers’ id columns to each other and expect a person match.
account_id_type is id, uuid, or email. match_method records how
the link was made (canonical, email_exact, email_local_part,
email_fuzzy, name_fuzzy). Identity rows are best-effort: service
accounts and shared mailboxes get identities too.
Worked tables
Linearissues (dataKind: changelog): one row per id. Order on
updatedat. Nested refs (assignee, team) stay on the issue row as
flattened or nested columns; they are not a second grain. Filter
WHERE id = ... for a single issue.
GitHub pull_requests (dataKind: changelog): one row per id.
Order on updated_at. Parent context copies full_name onto the PR so
you can filter by repo without joining repositories. Labels on GitHub
issues explode to providers.github.issues__labels, unique on label
id, with the parent issue id repeated.
Google users (dataKind: snapshot, identityDirectory: true): one
row per directory id (also the account id). primaryemail is the
person-email join column. Nested arrays (phones, external ids) are
separate tables with their own composite dedup keys.