Cohortum

Supported formats & limits

File formats, required columns, timestamp formats, scale, retention, and usage quota that Cohortum accepts for your event data.

This is the fine print: what your file needs before Cohortum can turn it into charts, how much you can load, and how long your data sticks around. Skim it before your first upload.

File formats

Cohortum reads CSV and Parquet, and each source sticks to one. You pick which in the first step of the upload wizard, and every load into that source has to match.

FormatExtensionNotes
CSV.csvNeeds a header row — the first line names your columns. Comma-delimited with standard double-quote quoting. Save as UTF-8.
Parquet.parquetRead with ClickHouse's native reader, about twice as fast as the default path on files with heavy JSON columns. Column names and types come from the file's schema. Internal compression (snappy, gzip, or zstd) is read transparently.

You pick the format on the first step of the upload wizard. Send CSV files uncompressed — plain .csv, not .csv.gz. For Parquet, keep whatever compression your export tool emits.

Prefer Parquet for large exports

If your warehouse or pipeline can emit Parquet, use it. It's smaller on the wire, carries its own column types, and imports faster than the same data as CSV.

Required columns

Every source needs three columns mapped, no matter the format. You assign them on the Mapping step of the upload wizard, in the Map to dropdown — and Cohortum auto-maps the obvious matches for you to review.

Map toWhat it holdsExpected typeExample
Event NameThe action a user tookTextVideo clicked, Page pagination clicked
User IDA stable per-user identifierTextuser_8f21, a3c9-…
TimestampWhen the event happenedDate/time (see below)2026-03-14T09:21:05Z

Any column you don't map falls to Ignore and gets skipped. The wizard reminds you: columns mapped as Ignore won't be imported.

Matching is case-sensitive

User IDs and event names are matched exactly, casing included. user_1 and User_1 are two different users; Signup and signup are two different events. Keep your casing consistent across every load into the same source.

Property columns as JSON

Past the three required columns, you can attach properties — the extra context you'll filter, split, and compare on later. Cohortum stores these as JSON-string columns, and there are two kinds:

  • Event Properties (JSON string) — attributes of a single event, like {"button":"play","position":3}.
  • User Properties (JSON string) — attributes of the person, like {"plan":"pro","country":"JP"}.

Each column holds one JSON object per row, written as a string. Point the Map to dropdown at the matching role, and the keys inside those objects show up as event and user properties across the analysis filter panel and the graph side-panels.

User properties are read as a snapshot per user, so the latest values win. Keep slow-changing traits (plan, country, acquisition channel) in User Properties and per-action details in Event Properties.

Accepted timestamp formats

Cohortum parses the Timestamp column leniently, so most common formats work as-is — no reformatting. It keeps values to millisecond precision.

StyleExample
ISO 8601 / RFC 33392026-03-14T09:21:05Z, 2026-03-14T09:21:05.482+00:00
Space-separated date-time2026-03-14 09:21:05
Date only2026-03-14
Unix epoch seconds1773480065
No offset means UTC

A value with no explicit UTC offset — a space-separated date-time, a date-only value, or any string without a Z or +HH:MM — is read as UTC. If your export uses a local timezone, add the offset (or convert to UTC) before you upload, or your events land in the wrong hour, day, and session.

Cohortum drops rows during import when a required column is unusable: a timestamp it can't parse, or a missing or empty Event Name or User ID. Dropped rows show up in the load history as Rows rejected, so you catch a bad file fast.

If you bring your own session identifier, map it on the Sessions step as a Session column — a numeric session id. Otherwise Cohortum builds sessions from a Session timeout (minutes) gap. See Prepare your data to shape all of this before you upload.

Event volume and upload size

A single source holds up to hundreds of millions of events. That's the technical ceiling — how far the engineering pushes. How much you can actually load comes down to your plan's usage quota below.

There's no fixed cap per file or per load. Files above 16 MB upload in parallel parts automatically, so a huge export won't stall on one slow transfer. You'll watch progress on the Upload step and a per-load status of pending → parsing → transforming → done in the source's Loads history.

Data retention

Retention is one window for the whole workspace, set under Settings → Workspace → Data retention in the Keep events for control. Anything older than the window gets deleted for good.

OptionMeaning
30 daysRolling 30-day window
90 daysRolling 90-day window
180 daysDefault
1 yearRolling 365-day window
UnlimitedNo age-based deletion
Deletion is permanent

Shorten the window and any events outside it are gone, with no undo. If you need a longer history, widen the window before you load older data. Only an Owner or Admin can change this setting.

Usage quota

Each workspace runs against a plan with an events allowance. Settings → Workspace → Usage shows your plan badge, an events used vs. limit count (say, 4.2M / 10M events), and the percentage you've burned through.

The quota is a soft limit — a running counter, not a hard cutoff. A load in progress won't get blocked or rejected when you cross the line, and ingestion won't quietly stop. Going over just means the Usage reading shows more than 100% until you bring it back down. As you near or pass the limit, prune old sources (or shorten your retention window) to free up room, or talk to your Cohortum contact about a bigger plan. There's no self-serve plan change in the app.