kimo
EngineeringSeries: Kimo Bridge

Cloud, hybrid or bridge: where should your analytics data live?

Put each data source where its strictest constraint says it should live: keep regulated or contract-bound production data at home and query it live through a bridge, sync high-volume or third-party SaaS data into the cloud for speed and history, and accept that most teams end up hybrid. The decision is made per source, using five questions — sensitivity, freshness, volume, query load and cost — not once for the whole company.

Arno Visser
Solutions architect10 min read7 sources

Every analytics project eventually hits the same meeting. Security wants to know where the data will be stored. Finance wants to know what it will cost. The data team wants dashboards that load in under two seconds. And legal has a contract clause that nobody has read in three years. The usual outcome is a blanket rule — "everything goes to the warehouse" or "nothing leaves our network" — that is wrong for half of the sources it covers.

As a solutions architect at Kimo, I spend most of my week in that meeting. This post is the framework I use to get out of it with a decision everyone can sign. It applies to any stack, but I will use Kimo’s two modes — Bridge (live query pushdown through Kimo Bridge) and Cloud (sync into Kimo’s managed cloud) — as the concrete options.

What are the three options, exactly?

Bridge mode vs Cloud modeBridge mode runs queries where the data lives and returns only results; Cloud mode keeps a synced copy next to the query engine.BRIDGE MODE · LIVE PUSHDOWNCLOUD MODE · MANAGED SYNCPostgreSQLYour databasestays putkimo-bridgeon your serverKimoplans the querySQL in · aggregates outStored on KimoNothing (optional short cache)FreshnessLive, on every querySpeedBound by your databaseBest forSensitive, regulated dataStripeYour sourcesDBs · SaaS APIsSyncscheduled · CDCKimo cloudencrypted · your regionStored on KimoEncrypted copy, your regionFreshnessPer schedule, e.g. 15 minSpeedSub-second, cachedBest forHistory, heavy dashboardsHybrid: choose per source, mix freely
Figure.Bridge mode runs queries where the data lives and returns only results; Cloud mode keeps a synced copy next to the query engine.

Scroll sideways to see the full diagram.

The five questions that decide where a source should live

Run every source through these questions in order. The first two can veto an option outright; the last three are trade-offs.

  1. Sensitivity and obligations. Does the source contain personal data, regulated data, or data a customer contract restricts? If yes, the burden is on anyone proposing a copy to justify it.
  2. Residency and transfers. Would a copy move data into another jurisdiction or to another processor? That triggers legal work before any engineering work.
  3. Freshness. How old can an answer be before it is wrong? Minutes, hours or a day?
  4. Volume and query shape. Do typical questions return small aggregates, or do analysts scan months of raw events interactively?
  5. Load and cost. Can the source absorb analytical queries, and what does each option cost in compute and data transfer?

1. Sensitivity: every copy is a new obligation

Under the GDPR, personal data must be "limited to what is necessary" for the purpose and kept in identifiable form no longer than necessary.1 A synced analytics copy is not forbidden by that principle, but it is a second system you have to justify, secure, include in retention schedules and purge on erasure requests. Live pushdown sidesteps much of that because aggregates — revenue by month, signups by channel — leave the database, while the rows behind them do not.

2. Residency: transfers and processors need paperwork

If a copy of personal data lands outside the European Economic Area, Chapter V of the GDPR applies: a transfer is only lawful if the conditions of that chapter are met by both controller and processor.2 The European Data Protection Board’s recommendations on supplementary measures describe how exporters should assess the destination and add safeguards where needed.3 And any vendor that stores the copy is a processor you must vet for "sufficient guarantees".4 None of this makes cloud sync impossible — Kimo’s cloud lets you pick a region — but it is the reason residency-sensitive sources often start in Bridge mode.

3. Freshness: live is free, fresh copies are not

With pushdown, every dashboard view reads the current state of the database. With sync, freshness equals the sync interval plus the time to process it. Going from daily to every 15 minutes is usually possible, but it multiplies sync jobs and source load. If a team cares about "right now" — open incidents, today’s signups, an investor asking about this month’s MRR mid-call — live wins.

4. Volume and query shape: where does the scan happen?

Pushdown is efficient when the database does the heavy lifting and returns a small result. It is inefficient when analysts repeatedly scan large histories on a transactional database that was not built for it. If your typical question is "weekly active accounts for the last two years, sliced six ways", a columnar copy in the cloud will be faster and will not slow your application down.

5. Load and cost: compute and egress

Live queries consume CPU on your side. Run them on a read replica — PostgreSQL standby servers accept read-only queries5 — and the primary is unaffected. On the cost side, remember data transfer pricing is asymmetric on major clouds: AWS, for example, charges nothing for inbound transfer but bills outbound transfer to the internet per service and region.6 A full sync of a large table out of your cloud account repeatedly pays that egress; a pushdown query returning a few hundred rows barely registers.

A decision matrix you can use today

If the source…Default toWhy
Holds personal, regulated or contract-restricted dataBridgeNo third-party copy to justify, secure or purge
Must stay in a specific country or your own infrastructureBridgeData never leaves; only query results cross the tunnel
Needs answers that are minutes old or lessBridgeEvery view reads current state
Is a SaaS API (ads, social, CRM, billing)CloudThe data already lives with a third party; APIs are slow and rate-limited
Is large event or log data scanned interactivelyCloudColumnar storage and caching beat repeated scans on an OLTP database
Needs history the source does not keep (snapshots, deleted rows)CloudOnly a synced copy can keep yesterday’s state
Is sensitive and largeBridge to a replica, or Cloud in your chosen regionDecide with security; pre-aggregate if possible
Rows are ordered by precedence: the first row that applies usually wins.

Worked example: a 60-person B2B SaaS company

Consider an illustrative company — call it Northwind Cloud — with an EU-hosted Postgres production database, a ClickHouse event store, Stripe billing, HubSpot, Google Ads and Search Console. Here is how the framework sorts it:

SourceModeReasoning
Postgres (accounts, users, usage)Bridge → read replicaPersonal data, EU residency clause in enterprise contracts, live usage questions
ClickHouse (product events, 2 years)BridgeAlready a fast columnar store (unlike an OLTP database), so pushdown performs well; no reason to duplicate billions of rows
StripeCloudThird-party API; needs MRR snapshots and history Stripe does not expose directly
HubSpot, Google Ads, Search ConsoleCloudRate-limited APIs; marketing wants trends across channels
Illustrative example. Your classification will differ.

The result is hybrid: the semantic layer joins live account data from the bridge with synced billing and marketing data, so the board deck and the marketing command center both read from one set of definitions, wherever the rows physically sit.

The sources in the worked example.

Common mistakes when choosing

  • Deciding once for the whole company. A blanket rule optimizes for the most sensitive source and punishes the rest, or the reverse.
  • Pointing live queries at the primary. Use a replica or a dedicated analytics role with statement timeouts.
  • Forgetting caches. A result cache is a copy too. In Kimo it is per source, has a TTL and can be set to zero; document what you chose.
  • Ignoring history. Bridge mode shows the database as it is now. If you need "what did this look like last quarter" and the source overwrites rows, you need snapshots — either in your own warehouse or in Cloud mode.
  • Treating network access as trust. An IP allowlist into your database is a standing path. Outbound-only, mutually authenticated connections with per-request authorization follow the zero-trust model instead.7

Write it down: the one-page source register

For each source, record

  • Owner and system of record
  • Data classes (personal, financial, regulated, public)
  • Contractual or regulatory constraints, with a link to the clause
  • Chosen mode (Bridge or Cloud), region, and cache TTL
  • Freshness requirement and sync schedule if any
  • Database role and exposed schemas
  • Date of next review

This register is what security reviewers and auditors actually want to see, and it makes the next change cheap: when a contract adds a residency clause, you flip one source from Cloud to Bridge instead of re-architecting. In Kimo, the Bridge page shows each source’s mode, and the connectors page shows sync schedules, so the register stays honest.

For the deeper architecture — tunnel design, caching, identity passthrough and how this compares to warehouse-first stacks — read our whitepaper Your Data, Your Rules. If you are ready to try it, start with Install Kimo Bridge with Docker.

Frequently asked questions

Is a bridge slower than a cloud warehouse?

For small aggregates on an indexed database, live pushdown is usually fast enough for interactive dashboards. For repeated scans over large histories, a synced columnar copy is typically faster. That is why the choice is per source.

Does Bridge mode make me GDPR compliant?

No architecture does on its own. Bridge mode reduces the personal data a vendor stores, which simplifies minimisation, retention and transfer analysis, but you still need a lawful basis and appropriate agreements.

Can I switch a source from Cloud to Bridge later?

Yes. In Kimo the mode is a per-source setting. Dashboards keep working because they query the semantic layer, not the storage location.

What does hybrid mean in practice?

Some sources are queried live through the bridge, others are synced to the cloud, and one semantic model joins them so every dashboard uses the same definitions.

Will live queries slow down my application?

They can if you point them at the primary. Use a read replica, a dedicated read-only role, statement timeouts and the bridge’s rate limits.

Sources

7 references
  1. Art. 5 GDPR: Principles relating to processing of personal data (opens in a new tab)
    gdpr-info.eu (Regulation (EU) 2016/679)2016gdpr-info.eu

    Data minimisation and storage limitation.

  2. Art. 44 GDPR: General principle for transfers (opens in a new tab)
    gdpr-info.eu (Regulation (EU) 2016/679)2016gdpr-info.eu

    Transfers to third countries only if Chapter V conditions are met.

  3. Art. 28 GDPR: Processor (opens in a new tab)
    gdpr-info.eu (Regulation (EU) 2016/679)2016gdpr-info.eu

    Controllers must use only processors providing sufficient guarantees.

  4. Hot Standby (opens in a new tab)
    PostgreSQL Documentationpostgresql.org

    Read-only queries on standby servers.

  5. Overview of Data Transfer Costs for Common Architectures (opens in a new tab)
    AWS Architecture Blog2021aws.amazon.com

    Inbound transfer is free; outbound to the internet is charged per service and region.

  6. SP 800-207: Zero Trust Architecture (opens in a new tab)
    NIST2020csrc.nist.gov

    No implicit trust based on network location.

External sources were accessed at the time of writing. Kimo product details, customers and figures in examples are illustrative unless a source is cited.

  • #Kimo Bridge
  • #Data residency
  • #Architecture
Found this useful? Pass it on.
Written by
Arno Visser
Solutions architect at Kimo · 1 article

Writes about Kimo Bridge, Data residency, Architecture.

Kimo people and customers mentioned are illustrative; example charts use simulated data unless a source is cited.

Put it to work

Go deeper

Whitepaper

Your Data, Your Rules

The hybrid analytics architecture behind Kimo Bridge: live query pushdown, optional cloud sync, and zero-trust by default.

22 pages
Live demo

Set up Kimo Bridge

Install the bridge, pick Bridge or Cloud mode per source, watch the audit log.

Simulated data · no sign-up

All resources
ArticleProduct
All

Introducing Kimo Bridge: your data stays home

A small package you install next to your database opens a private, outbound-only bridge to Kimo. Query live, store nothing — or sync to our cloud when you want to.

Théo Marchand
7 min read
Whitepaper
All

Your Data, Your Rules

The hybrid analytics architecture behind Kimo Bridge: live query pushdown, optional cloud sync, and zero-trust by default.

Arno Visser
22 pages
GuideIntermediate
All

The Kimo Bridge security model

What leaves your network, what never does, and how every query is authorized and audited.

Rhea Patel
10 min read

Your data officer is ready.

Connect a source — or install Kimo Bridge and keep data on your servers — then ask a question and get an answer you can audit.