Course overviewGoogle Search · Prepare text and combine files
Recognize duplicate records
Identify repeated event IDs without deleting distinct searches that share query text.
Step 1 of 3 · Learn
Define the duplicate key
.duplicated(subset=["id"]) marks repeated keys after their first occurrence. Use keep=False to flag every row in a repeated-key group. Matching query text alone does not mean two rows represent the same search.
index
id
query
0
1
tea
1
2
tea
2
1
tea
↓
index
id
query
repeated
0
1
tea
true
1
2
tea
false
2
1
tea
true
ID 1 repeats. ID 2 is a different event despite identical query text.
Choose which occurrence survives
.drop_duplicates(subset=["id"], keep="first") keeps the first occurrence; keep="last" keeps the last in current row order. Neither option checks timestamps. Without subset, Pandas compares all columns, so conflicting versions may both remain.
index
id
value
0
1
10
1
2
20
2
1
15
↓
index
id
value
1
2
20
2
1
15
Keep the last occurrence of each ID, preserving the surviving rows' order.
▷ Your turn
Return every row of imported_searches with columns search_id, query, response_ms, and repeated_id. Set repeated_id to True for every occurrence of an ID appearing more than once, otherwise False. Preserve import order and save the DataFrame as result.
imported_searches is a six-row working copy based on four searches. Two rows repeat earlier IDs; the last row changes response_ms by 25. Two different IDs share a query. Row order represents import order, not event time. The original searches is unchanged.