Supported formats & limits
File formats, required columns, timestamp formats, scale, retention, and usage quota that Cohortum accepts for your event data.
This is the fine print: what your file needs before Cohortum can turn it into charts, how much you can load, and how long your data sticks around. Skim it before your first upload.
File formats
Cohortum reads CSV and Parquet, and each source sticks to one. You pick which in the first step of the upload wizard, and every load into that source has to match.
| Format | Extension | Notes |
|---|---|---|
| CSV | .csv | Needs a header row — the first line names your columns. Comma-delimited with standard double-quote quoting. Save as UTF-8. |
| Parquet | .parquet | Read with ClickHouse's native reader, about twice as fast as the default path on files with heavy JSON columns. Column names and types come from the file's schema. Internal compression (snappy, gzip, or zstd) is read transparently. |
You pick the format on the first step of the upload wizard. Send CSV files uncompressed — plain .csv, not .csv.gz. For Parquet, keep whatever compression your export tool emits.
If your warehouse or pipeline can emit Parquet, use it. It's smaller on the wire, carries its own column types, and imports faster than the same data as CSV.
Required columns
Every source needs three columns mapped, no matter the format. You assign them on the Mapping step of the upload wizard, in the Map to dropdown — and Cohortum auto-maps the obvious matches for you to review.
| Map to | What it holds | Expected type | Example |
|---|---|---|---|
| Event Name | The action a user took | Text | Video clicked, Page pagination clicked |
| User ID | A stable per-user identifier | Text | user_8f21, a3c9-… |
| Timestamp | When the event happened | Date/time (see below) | 2026-03-14T09:21:05Z |
Any column you don't map falls to Ignore and gets skipped. The wizard reminds you: columns mapped as Ignore won't be imported.
User IDs and event names are matched exactly, casing included. user_1 and User_1 are two different users; Signup and signup are two different events. Keep your casing consistent across every load into the same source.
Property columns as JSON
Past the three required columns, you can attach properties — the extra context you'll filter, split, and compare on later. Cohortum stores these as JSON-string columns, and there are two kinds:
- Event Properties (JSON string) — attributes of a single event, like
{"button":"play","position":3}. - User Properties (JSON string) — attributes of the person, like
{"plan":"pro","country":"JP"}.
Each column holds one JSON object per row, written as a string. Point the Map to dropdown at the matching role, and the keys inside those objects show up as event and user properties across the analysis filter panel and the graph side-panels.
User properties are read as a snapshot per user, so the latest values win. Keep slow-changing traits (plan, country, acquisition channel) in User Properties and per-action details in Event Properties.
Accepted timestamp formats
Cohortum parses the Timestamp column leniently, so most common formats work as-is — no reformatting. It keeps values to millisecond precision.
| Style | Example |
|---|---|
| ISO 8601 / RFC 3339 | 2026-03-14T09:21:05Z, 2026-03-14T09:21:05.482+00:00 |
| Space-separated date-time | 2026-03-14 09:21:05 |
| Date only | 2026-03-14 |
| Unix epoch seconds | 1773480065 |
A value with no explicit UTC offset — a space-separated date-time, a date-only value, or any string without a Z or +HH:MM — is read as UTC. If your export uses a local timezone, add the offset (or convert to UTC) before you upload, or your events land in the wrong hour, day, and session.
Cohortum drops rows during import when a required column is unusable: a timestamp it can't parse, or a missing or empty Event Name or User ID. Dropped rows show up in the load history as Rows rejected, so you catch a bad file fast.
If you bring your own session identifier, map it on the Sessions step as a Session column — a numeric session id. Otherwise Cohortum builds sessions from a Session timeout (minutes) gap. See Prepare your data to shape all of this before you upload.
Event volume and upload size
A single source holds up to hundreds of millions of events. That's the technical ceiling — how far the engineering pushes. How much you can actually load comes down to your plan's usage quota below.
There's no fixed cap per file or per load. Files above 16 MB upload in parallel parts automatically, so a huge export won't stall on one slow transfer. You'll watch progress on the Upload step and a per-load status of pending → parsing → transforming → done in the source's Loads history.
Data retention
Retention is one window for the whole workspace, set under Settings → Workspace → Data retention in the Keep events for control. Anything older than the window gets deleted for good.
| Option | Meaning |
|---|---|
| 30 days | Rolling 30-day window |
| 90 days | Rolling 90-day window |
| 180 days | Default |
| 1 year | Rolling 365-day window |
| Unlimited | No age-based deletion |
Shorten the window and any events outside it are gone, with no undo. If you need a longer history, widen the window before you load older data. Only an Owner or Admin can change this setting.
Usage quota
Each workspace runs against a plan with an events allowance. Settings → Workspace → Usage shows your plan badge, an events used vs. limit count (say, 4.2M / 10M events), and the percentage you've burned through.
The quota is a soft limit — a running counter, not a hard cutoff. A load in progress won't get blocked or rejected when you cross the line, and ingestion won't quietly stop. Going over just means the Usage reading shows more than 100% until you bring it back down. As you near or pass the limit, prune old sources (or shorten your retention window) to free up room, or talk to your Cohortum contact about a bigger plan. There's no self-serve plan change in the app.