Prepare your data
Shape your event log for Cohortum with three required columns — User ID, Event Name, Timestamp — plus optional properties and sessions.
Cohortum builds its charts straight from your raw event log. Shape the log before you upload it, and mapping takes a few clicks.
The event-log data model
Everything in Cohortum comes down to one row per event. Each row says a user did something at a moment in time. Steps, Journey, Transitions, and Clusters all read from those rows, so a clean log means a clean analysis.
You don't have to ship every raw signal untouched. It often helps to merge events that mean the same thing into a single label — collapsing near-duplicate variants of one action into one name — to cut noise, and to drop events that carry little signal, like debug pings, heartbeats, or ultra-frequent UI ticks. A tighter set of meaningful events makes sessions, paths, and segments far easier to read.
Required columns
Every source needs three columns. During upload, the Mapping step asks you to point each column in your file at one of these roles.
| Your column (example) | Maps to | What it holds |
|---|---|---|
user_id | User ID | A stable identifier for the person — the same value across all of their events and sessions. |
event | Event Name | What happened, as a short human-readable label (e.g. Video clicked). |
timestamp | Timestamp | When the event occurred. |
Your column names don't have to match the role names. Cohortum guesses the likely mapping — the example names above are the kind it picks up — and you confirm or change each one. All it cares about is that the three roles get filled.
Your file also needs a header row, so Cohortum has column names to show in the mapping table. Save it as UTF-8 so non-ASCII event names and property values survive the trip — a country of JP, or an event written in another language.
Example CSV
Here's a small, well-formed event log. The first three columns are the required roles; the last two are optional property columns, covered below.
user_id,event,timestamp,event_properties,user_properties
u_10432,Video clicked,2026-07-01T09:14:22Z,"{""video_id"":""v_88"",""position"":3}","{""plan"":""pro"",""country"":""JP""}"
u_10432,Page pagination clicked,2026-07-01T09:15:04Z,"{""page"":2}","{""plan"":""pro"",""country"":""JP""}"
u_20981,Video clicked,2026-07-01T09:15:41Z,"{""video_id"":""v_12"",""position"":1}","{""plan"":""free"",""country"":""US""}"
u_20981,Checkout completed,2026-07-01T09:22:10Z,"{""amount"":49.00}","{""plan"":""free"",""country"":""US""}"
u_33107,Video clicked,2026-07-01T09:31:55Z,"{""video_id"":""v_88"",""position"":2}","{""plan"":""pro"",""country"":""DE""}"Both rows for u_10432 carry the same user and property values. That's how Cohortum threads one person's actions into an ordered path.
Optional columns
Add properties to your events and you unlock filtering, comparison overlays, and the numeric insights on every tab.
Event properties and user properties
Send both as JSON string columns — one column whose cells each hold a JSON object.
- Event properties (JSON) describe the event:
{"video_id":"v_88","position":3}. They change row to row. - User properties (JSON) describe the person:
{"plan":"pro","country":"JP"}. They usually repeat across a user's rows.
In the Mapping step, point each JSON column at User Properties (JSON) or Event Properties (JSON). The keys inside turn into filterable, comparable event and user properties across the workspace.
You get one Event Properties column and one User Properties column per source — the Map-to dropdown offers exactly those two property roles, not a slot per attribute. So if your export drops each property in its own flat column (video_id, position, plan, country as separate columns — what spreadsheets, CRMs, and most analytics tools spit out), fold them into a single JSON column before upload. Pack the event-level attributes into one JSON object per row and the user-level attributes into another. A spreadsheet formula that concatenates {"key":value,…}, or a to_json/json_object step in your export pipeline, handles it in one pass. Keep the values flat — strings and numbers. Nested objects and arrays won't become their own filterable keys, and a given key should hold the same value type in every row.
An explicit session column
Out of the box, Cohortum auto-generates sessions: it starts a fresh one after a stretch of inactivity. You set how long that gap is with a Session timeout (minutes) on the Sessions step, so you pick the window at upload time. Already have a session identifier in your log? Use it instead — include a session column and select it in the Sessions step of the upload wizard. You choose between Auto-generate sessions and a Session column right there.
The "Session" concept drives Sessions Analysis and every session-based filter. If your source system already cuts sessions the way you think about them, mapping an explicit session column keeps Cohortum in step with that definition. See Sessions for how boundaries are formed and used.
Accepted file formats
Cohortum reads CSV and Parquet. You pick the type at the start of the upload wizard. Reach for Parquet on large, property-heavy logs; CSV is simplest for exports from spreadsheets and most tools.
Timestamp format
Give Cohortum an unambiguous, machine-readable timestamp. ISO 8601 (for example 2026-07-01T09:14:22Z) is the safest bet — it sorts correctly and can't be misread. Include the time of day, not just the date. Cohortum orders each user's events by timestamp to rebuild their path, so second-level precision keeps the sequence right.
Ground rules for a clean log
Use one stable label per action. Video clicked, video_clicked, and Clicked video land as three separate events and split your paths. Pick a naming style and stick to it everywhere. If some drift slips through, you can still rename events in the Events catalog after import — but a consistent log up front saves the cleanup.
Cohortum needs one unique, stable identifier per person that follows them across sessions and devices — an account ID or a persistent anonymous ID. Don't put sensitive data here — no emails, names, or other PII: the User ID is only a key for grouping events, never a contact detail. The ideal value is a hash of your production user ID (for example a SHA-256 of the account ID) — stable, unique, and impossible to read back. Skip per-session or per-visit identifiers, or every visit looks like a brand-new user and your cohorts fall apart.
- Session IDs used as the User ID — the top cause of inflated user counts and broken journeys.
- Inconsistent event names — drift in casing and wording splits one action into many.
- Ambiguous timestamps — locale-dependent formats like
07/01/26can be misread; prefer ISO 8601. - Empty required fields — rows missing a user, event, or timestamp can't be placed in a path.
Follow these rules and mapping takes a moment — Cohortum auto-maps the columns and you just review the result. Not sure your file will import cleanly? Check the size and type limits first.