Users
The user id column ties every event to a person — it drives per-user sequences, the "N users" count, and every cohort you build.
Every event in your clickstream came from someone. The user id column is how Cohortum knows which events came from the same person — and almost everything you see in the product is built on that one column.
What the user id does
When you upload an event log, you map three required columns: Event Name, User ID, and Timestamp (see Prepare your data). The user id is what turns a pile of loose events into a story.
Here's what it powers:
- Per-user sequences. Cohortum grabs all of a person's events, sorts them by time, and lays them out as one path — "Home viewed" → "Video clicked" → "Page pagination clicked". Now every chart knows what someone did right before and right after each step.
- The unique-users count. The users figure in the analysis toolbar (say, "8.4k users") counts distinct user ids that match your filters — people, not events or rows.
- Cohorts. A User Cohort is just a set of user ids. Drill into a step, upload a CSV list, or save a filter — either way you end up with a list of people, identified by this column.
Drop-off in Steps, edges in Transitions, paths in Journey, and segments in Clusters all count users, not raw events. Get the user id right and every tab shows you real people. Get it wrong and every number is off the same way.
What makes a good user id
The same person should carry the same id everywhere — across sessions, days, and devices. That's the whole game. When the id is stable and consistent, Cohortum recognizes a returning user instead of treating every visit as a stranger.
Good choices:
- A signed-in account id or internal user id from your database.
- A durable anonymous id — a first-party cookie or device id — that sticks around between visits.
Keep it non-sensitive. A user id is only a key for grouping events, never a contact detail — so don't map emails, names, or other PII into it. The ideal value is a hash of your production user id (for example a SHA-256 of the account id): stable and unique per person, but impossible to read back to a real identity.
Ids to avoid
Steer clear of anything that changes more often than the person does. When the id resets on every visit, Cohortum sees a crowd of one-event users: your unique-user count balloons and your multi-step paths fall apart.
| Avoid | Why it breaks analysis |
|---|---|
| A per-session id | Each session looks like a different person; return behavior disappears. |
| A per-request / per-event id | Every event is its own "user"; sequences never form. |
A shared or placeholder id (e.g. all 0 or blank) | Many people collapse into one; paths become nonsense. |
A quick red flag: if the users count in the analysis toolbar is close to your total event count, or if Steps and Journey show almost nothing past the first step, your User ID mapping is probably wrong.
Cohortum tracks sessions separately — one window of activity. That session key belongs in your source's session configuration, not in the User ID mapping. Mapping a session id as the user id is the single most common data mistake. See Sessions for how sessions work.
Identified vs. anonymous users
Both work. Map a logged-in account id and you're analyzing known users. Map an anonymous device or cookie id and you're analyzing anonymous visitors — that's just as valid.
What you can't do is mix the two inconsistently — the same person showing up under an anonymous id before login and an account id after. Now one person splits into two users, and their journey gets cut in half right at the login. If you can stitch the anonymous id to the account id before export, do it there. Then Cohortum gets one clean id per person.
Each analysis covers a single source, so ids don't need to line up across separate sources for a normal analysis. But if you want one person's web and mobile activity in a single journey, those events have to share one id inside the same source. Combine the logs before you upload — don't load two sources and hope to join them later.
Capitalization, whitespace, and numeric-vs-string formatting each make a different id: U_123 and u_123 are two people, and 007 turns into 7 the moment Excel strips the leading zero. This bites hardest when you upload a cohort CSV — if the ids in your list don't match the event log exactly, the build finishes with zero matching users.
If you already loaded data with the wrong id
You set a source's column mapping once, in the wizard's Mapping step, and it locks the moment the source finishes loading — there's no in-place remap. So if a loaded source used the wrong User ID column, don't try to edit it. Create a new source with the correct mapping and load your file again. See Sources & loads for how mapping locks, and Prepare your data to get it right the first time.