Push data to Ronja
Pushing means your own code sends data to Ronja, on its own schedule: a nightly export, a script that calls a supplier’s API, an app that records events. This page is the playbook: which path to pick, and how to keep each run fast, safe to repeat and cheap on rebuilds. The endpoint-by-endpoint walkthrough, with every request and response, is Send data via the API.
Pick the shortest path first
Section titled “Pick the shortest path first”| Where the data is | What to use | Why |
|---|---|---|
| In your own database: PostgreSQL, MySQL, SQL Server or BigQuery | A connection (Admin), not push code. See Send data via the API | Ronja syncs the tables you choose on a schedule and keeps track of how far she got. There is no script for you to keep running |
| In a file you have once | Upload files | No code at all |
| Rows your own code produces: API results, readings, a summary you compute | A Dynamic table, row push | The write is synchronous: when it returns, the rows are queryable. No build step |
| Parquet files you already have: a warehouse export, a nightly dump | An Integration table, file push, then Build | Files accumulate by filename and are built together |
Send few, large writes
Section titled “Send few, large writes”Every write is an update to the table. It rebuilds the Derived tables built from it and fires every Automation watching it, and an append adds a data file. So the cheapest run is the one with the fewest writes:
- One write per table per run. Collect the run’s rows, then send them in one request, not one request per record. A script that writes every few seconds rebuilds everything downstream every few seconds and fires its Automations just as often.
- Send Parquet, not JSON, once the data is large. A JSON write takes at most 50,000 rows or 8 MiB. A Parquet upload to the same table takes up to 1 GiB with no row limit, and the same rows are a small fraction of the bytes. With pandas it is one line:
df.to_parquet(buf). - On JSON, send values with their real types. Numbers as numbers (
"qty": 3, not"qty": "3") and trimmed strings. A quoted number on a column of numbers is a type change (on a table still on the older reading of JSON, see Push rows). See When a column changes type.
Make every run safe to repeat
Section titled “Make every run safe to repeat”A run that fails halfway gets run again, by you or by a scheduler, so make a repeat harmless:
- Refresh with
replace. If each run has the whole current dataset, send it with"writeMode": "replace". Sending it twice leaves the same table. Keepappendfor records that are genuinely new, and know that anappendthat timed out may still have landed, so retrying it can write the batch twice. - Name pushed files after what they hold. On an Integration table, re-uploading a filename replaces that file. A file named by its date (
orders_2026-10-03.parquet) is overwritten by a retry rather than added a second time. - Retry the answers that say to. A 409 on a write means the table is busy, with another write or with Ronja tidying its data files: nothing was written, so wait a moment and send it again. A 503 on the release below is safe to retry too. A server error (500, 502, 504) or a dropped connection is safe to resend only for a call that is safe to repeat: a
replacewrite or the release, not anappend. - Store the table ID. Keep the
table-…ID from when you created the table rather than looking the table up by name on every run.
Write several tables as one change
Section titled “Write several tables as one change”When one run writes several tables that a Derived table reads together, such as orders and customers joined in a report, each write would rebuild the report on its own. The first rebuild joins the new orders with the old customers. Instead, hold back each table’s rebuild and release them all together:
- Send every write with
?cascade=defer. - Check that each response says
"cascade": "deferred". - After the last write, release them in one call:
POST /api/v2/feature/model/cascadewith{"tableIds": [...]}.
The report then rebuilds once, over all of the new data. The full rules are in Send several tables at once. Two of them shape how you write the script:
- Send everything first, then release straight away. A table nobody releases is released on its own about 15–20 minutes after its first held write, so a run that takes longer than that has its first tables released partway through.
- If a write fails partway, retry it before you release. Release only once every table has landed. A hold cannot be cancelled: if you give up, the tables you already wrote are released after 15–20 minutes anyway, and the Derived tables rebuild over the part of the set that landed. Fix the failure and run again soon.
A complete nightly push
Section titled “A complete nightly push”This script refreshes two Dynamic tables as one change. It sends each as Parquet with replace, so every call in it is safe to repeat, and it retries the busy and server-error answers. Give the client timeout room for your largest write: a write still running when the client gives up keeps landing, and the retries meet a 409 until it finishes.
import io, os, timeimport pandas as pdimport requests
BASE = os.environ["RONJA_BASE_URL"]S = requests.Session()S.headers["Authorization"] = f"Bearer {os.environ['RONJA_API_TOKEN']}"
TABLES = {"orders": "table-a1b2c3d4e5", "customers": "table-f6g7h8i9j0"}RETRY = {409, 500, 502, 503, 504}
def call(method, path, **kw): # Every call below is safe to repeat, so a failure that may be passing is retried. for attempt in range(6): if attempt: time.sleep(2 ** attempt) # 2, 4, 8, 16, 32 seconds try: r = S.request(method, BASE + path, timeout=900, **kw) except (requests.ConnectionError, requests.Timeout): continue if r.status_code not in RETRY: r.raise_for_status() # any other error is a real refusal: stop return r.json() raise RuntimeError(f"{method} {path} kept failing")
def push(name, df): buf = io.BytesIO() df.to_parquet(buf, index=False) out = call("POST", f"/api/v2/feature/model/{TABLES[name]}/rows/parquet", params={"writeMode": "replace", "cascade": "defer"}, headers={"Content-Type": "application/octet-stream"}, data=buf.getvalue()) print(name, out.get("cascade")) # "deferred": held for the release below
push("orders", load_orders()) # your own extractspush("customers", load_customers())
print(call("POST", "/api/v2/feature/model/cascade", json={"tableIds": list(TABLES.values())}))# -> {"released": ["table-a1b2c3d4e5", "table-f6g7h8i9j0"], "notPending": []}If push("customers", …) keeps failing, the script stops before the release. The orders write stays held until the automatic release, so rerun the script once the failure is fixed.
For an Integration table the loop is the same with files: upload each table’s files, start every Build with ?cascade=defer, wait for each build to finish, then release. A table that is still building is refused with a 409 and nothing is released, so wait for every build first. See Send several tables at once.