Pandas · Find patterns in text
Google Search Analytics
Course overviewGoogle Search · Prepare text and combine files

Find patterns in text

Match literal text safely and extract a defined URL pattern.

Step 1 of 3 · Learn

Match literal text

.str.contains(text, regex=False, na=False) produces a Boolean mask. regex=False treats punctuation literally; na=False makes missing input a non-match. Add case=False for case-insensitive matching.

indextext
0a.b
1axb
2NA
indextext
0a.b
A literal dot matches a dot, not any character.

Extract a specific pattern

.str.extract(pattern, expand=False) returns a Series for one capture group. In r"^https://([^/]+)", ^ anchors the start; [^/]+ matches one or more non-slash characters; parentheses capture that text. Non-matches stay missing.

indexurl
0https://example.org/a
1https://sample.net/b
2http://example.org/c
indexurlhost
0https://example.org/aexample.org
1https://sample.net/bsample.net
2http://example.org/cNA
Capture the text after the HTTPS prefix and before the next slash.

Match versus extract

Contains selects matching rows; extract keeps the row positions and returns captured values. This small regex fits the course URLs. It is not a general URL parser for credentials, ports, or other URL formats.

▷ Your turn

From searches, return search_id and query for rows containing the literal text "pandas", ignoring case. Missing queries must not match. Preserve the original query text and source row order. Save the DataFrame as result.

LANGUAGEPython · Pandas

Loading Python and Pandas…

Run the code to see DataFrame results here.