Course overviewGoogle Search · Prepare text and combine files
Find patterns in text
Match literal text safely and extract a defined URL pattern.
Step 1 of 3 · Learn
Match literal text
.str.contains(text, regex=False, na=False) produces a Boolean mask. regex=False treats punctuation literally; na=False makes missing input a non-match. Add case=False for case-insensitive matching.
index
text
0
a.b
1
axb
2
NA
↓
index
text
0
a.b
A literal dot matches a dot, not any character.
Extract a specific pattern
.str.extract(pattern, expand=False) returns a Series for one capture group. In r"^https://([^/]+)", ^ anchors the start; [^/]+ matches one or more non-slash characters; parentheses capture that text. Non-matches stay missing.
index
url
0
https://example.org/a
1
https://sample.net/b
2
http://example.org/c
↓
index
url
host
0
https://example.org/a
example.org
1
https://sample.net/b
sample.net
2
http://example.org/c
NA
Capture the text after the HTTPS prefix and before the next slash.
Match versus extract
Contains selects matching rows; extract keeps the row positions and returns captured values. This small regex fits the course URLs. It is not a general URL parser for credentials, ports, or other URL formats.
▷ Your turn
From searches, return search_id and query for rows containing the literal text "pandas", ignoring case. Missing queries must not match. Preserve the original query text and source row order. Save the DataFrame as result.