Skip to main content

Pipelines and scheduling

A pipeline is a sync job. It reads from one source connection and writes to one destination connection, and it handles schema normalization, incremental loading and type mapping for you. Each pipeline writes to a dataset of its own, so pipelines never collide in the destination.

If you have not made one yet, start with Getting started.

Running a pipeline​

Choose Run now on a pipeline to start a one-off sync. Manual runs and scheduled runs are independent, so running a pipeline by hand does not change its schedule.

Only one run per pipeline can be active at a time. Run now is disabled while a run is pending or running.

Run statuses​

A run moves through these states:

StatusMeaning
pendingThe run is queued and waiting to start. This is usually near-instant.
runningKoraBridge is fetching data from the source and writing it to the destination. The pipeline page refreshes every few seconds while a run is active.
successAll selected streams were fetched and loaded. Row counts and duration are recorded.
failedThe run hit an error. The run history shows a failure summary and the error message.
cancelledThe run was stopped before it finished. No further data is written for that run.

Reading run history​

Each run in the history shows:

  • New rows and rows processed. New rows are records that were newly fetched or changed. Rows processed is the total written in this run, including reference data that is snapshotted again. Streams that always reload in full show only a processed count.
  • Duration. The time from start to completion. Long durations point to large data volumes or slow destination writes.
  • Error message. On failed runs, a plain-language summary with suggested next steps comes first. Expand Technical details for the raw provider or engine error when you need it for support.
  • Load ID. A unique identifier for the load. Use it to find the matching rows in the load metadata in your destination, and to tell which output files belong to a run.

Stopping a run​

Choose Stop on the active run to ask KoraBridge to cancel it. If the run finishes on its own before the cancellation is processed, the stop does nothing and the run keeps its real outcome.

Schedules​

A schedule runs a pipeline automatically on a recurring basis.

  1. Open Schedules from the sidebar and create a schedule.
  2. Select the pipeline to automate and choose a type, Cron or Interval.
  3. Set the timing. For Cron, enter a cron expression. For Interval, enter the number of seconds between runs.
  4. Schedules start enabled. You can pause or delete them at any time.

A pipeline can have more than one schedule, for example an hourly GPS sync and a daily wellness sync.

Minimum interval​

Runs cannot be scheduled more often than once per hour. Cron and interval schedules are both checked against this limit. Sports vendor APIs are rate limited and their data arrives per session, match or day, so faster polling mostly fetches rows that have not changed.

Common cron expressions​

CadenceExpression
Every hour0 * * * *
Every 6 hours0 */6 * * *
Daily at midnight0 0 * * *

When a run fails​

  1. Open the pipeline page and read the failure summary and suggested next steps under the failed run.
  2. If the error mentions credentials or authentication, go to Connections and test the source or destination again.
  3. Expand Technical details only when you need the raw provider diagnostics for support.
  4. If the destination reports a schema error, check that the destination user is allowed to create and alter tables.
  5. Fix the cause, then choose Run now to retry.
  6. If a scheduled pipeline keeps failing, pause its schedule on the Schedules page while you investigate.

What happens to data when a run fails​

  • If a run fails while reading from the source, nothing was loaded and the incremental position does not advance. The next run picks up from the same point.
  • If it fails while writing to the destination, files or rows from jobs that already finished can remain. Only loads recorded in the _dlt_loads table are complete, so ignore files whose load ID is not listed there.
  • The next run always starts from the last complete position.

Incremental loading and full reloads​

KoraBridge keeps incremental state between runs, so a run fetches only what changed since the last complete one. There is no self-serve full reload yet.

Data layout in file destinations​

For Amazon S3, Azure Blob Storage and Azure Data Lake, each pipeline writes into its own dataset folder named t_<tenant prefix>_<pipeline id>. The name combines a short organization prefix with the full pipeline identifier.

  • Parquet files. Each run writes one or more .parquet files per table, named <load_id>.<chunk>.parquet. Runs append new files and old files are not deleted. Use the Load ID from the run history to find the files for one run.
  • Child tables. When a record contains a nested array, such as the squad IDs inside an athlete record, KoraBridge unpacks it into a child table named <parent>__<child>, with a double underscore. Join it back to the parent on the _dlt_root_id and _dlt_id columns.
  • Load metadata. The _dlt_loads and _dlt_pipeline_state tables hold load metadata and state. They contain no source data.

For example, in Spark:

df = spark.read.parquet("s3://<bucket>/<dataset>/athlete_users/")
df_squads = spark.read.parquet("s3://<bucket>/<dataset>/athlete_users__squad_ids/")
df.join(df_squads, df["_dlt_id"] == df_squads["_dlt_root_id"])

FAQ​

What is a pipeline exactly? A sync job that reads from one source connection and writes to one destination connection.

Can I run a pipeline and also schedule it? Yes. Manual and scheduled runs are independent.

How often can a schedule run? At most once per hour.

Does testing an S3 destination prove I can write to the bucket? Yes, in both access-key and IAM role mode. The test writes and deletes a small object, so a pass means the pipeline can write there. S3-compatible endpoints such as MinIO or Cloudflare R2 are tested the same way.