Pandas · Recognize duplicate records
Google Search Analytics
Course overviewGoogle Search · Prepare text and combine files

Recognize duplicate records

Identify repeated event IDs without deleting distinct searches that share query text.

Step 1 of 3 · Learn

Define the duplicate key

.duplicated(subset=["id"]) marks repeated keys after their first occurrence. Use keep=False to flag every row in a repeated-key group. Matching query text alone does not mean two rows represent the same search.

indexidquery
01tea
12tea
21tea
indexidqueryrepeated
01teatrue
12teafalse
21teatrue
ID 1 repeats. ID 2 is a different event despite identical query text.

Choose which occurrence survives

.drop_duplicates(subset=["id"], keep="first") keeps the first occurrence; keep="last" keeps the last in current row order. Neither option checks timestamps. Without subset, Pandas compares all columns, so conflicting versions may both remain.

indexidvalue
0110
1220
2115
indexidvalue
1220
2115
Keep the last occurrence of each ID, preserving the surviving rows' order.
▷ Your turn

Return every row of imported_searches with columns search_id, query, response_ms, and repeated_id. Set repeated_id to True for every occurrence of an ID appearing more than once, otherwise False. Preserve import order and save the DataFrame as result.

imported_searches is a six-row working copy based on four searches. Two rows repeat earlier IDs; the last row changes response_ms by 25. Two different IDs share a query. Row order represents import order, not event time. The original searches is unchanged.

LANGUAGEPython · Pandas

Loading Python and Pandas…

Run the code to see DataFrame results here.