How to use regex in BS4 `find_all` to return matched items with pattern priority?

Question

I have the following regex expression:

import re

re.compile('|'.join([pattern1, pattern2, pattern3]))

I would like it to work in the following way:

Try to match only pattern1; if matched - stop; else - proceed.
Try to match only pattern2; if matched - stop; else - proceed.
Try to match only pattern3; stop.

However currently it matches all of them.

I found this Q/A, which I thought answers my question, but adding flags=re.I does not fix my issue, since my result does not change.

How is this possible (if at all)?

A reproducible example:

from bs4 import BeautifulSoup

xml_doc = """
    
    """

soup = BeautifulSoup(xml_doc, "xml")

# This gives 11 vales.
len(soup.find_all(re.compile('|'.join([
    r'^m[0-9]_commodity_group$',r'^m[0-9]_region_group$',r'^m[0-9]_attribute_group$'
]), flags=re.I)))

# This gives 1 value <-- It's what I want, but I want to achieve it with the regex from above (which would work for other texts)
len(soup.find_all(re.compile('|'.join([
    r'^m[0-9]_commodity_group$'
]), flags=re.I)))

# This gives 10 values, which in this example I'd like to be ignored, since the first regex already gave results.
len(soup.find_all(re.compile('|'.join([
    r'^m[0-9]_attribute_group$'
]), flags=re.I)))

Talon · Accepted Answer

You could restructure your search:

patterns = [r'^m[0-9]_commodity_group$',r'^m[0-9]_region_group$',r'^m[0-9]_attribute_group$']
for pattern in patterns:
    result = soup.find_all(re.compile(pattern, flags=re.I))
    if result:
        break  # Stop after the first time you found a match
else:
    result = None  # When there never was a match

That might be more reabable than regex magic. If you will be executing this a lot, you might want to pre-compile your regexes once instead of at every loop iteration.

How to use regex in BS4 `find_all` to return matched items with pattern priority?

Answers (2)

Related Questions