Iterating through a table of rows with beautiful soup in python

Question

I'm trying to parse through a table of rows using beautiful soup and save values of each row in a dict.

One hiccup is the structure of the table has some rows as the section headers. So for any row with the class 'header' I want to define a variable called "section". Here's what I have, but it's not working because it's saying ['class'] TypeError: string indices must be integers

Here's what I have:

for i in credits.contents:
    if i['class'] == 'header':
        section = i.contents
        DATA_SET[section] = {}
    else:
        DATA_SET[section]['data_point_1'] = i.find('td', {'class' : 'data_point_1'}).find('p').contents
        DATA_SET[section]['data_point_2'] = i.find('td', {'class' : 'data_point_2'}).find('p').contents
        DATA_SET[section]['data_point_3'] = i.find('td', {'class' : 'data_point_3'}).find('p').contents

Example of data:


    
        HEADER NAME
    
    
        DATA
        DATA
        DATA
    
    
        DATA
        DATA
        DATA
    
    
        DATA
        DATA
        DATA
    
    
        HEADER NAME
    
    
        DATA
        DATA
        DATA
    
    
        DATA
        DATA
        DATA
    
    
        DATA
        DATA
        DATA

daedalus · Accepted Answer

Here is one solution, with a slight adaptation of your example data so that the result is clearer:

from BeautifulSoup import BeautifulSoup
from pprint import pprint

html = '''
    
        HEADER 1
    
    
        DATA11
        DATA12
        DATA12
    
    
        DATA21
        DATA22
        DATA23
    
    
        DATA31
        DATA32
        DATA33
    
    
        HEADER 2
    
    
        DATA11
        DATA12
        DATA13
    
    
        DATA21
        DATA22
        DATA23
    
    
        DATA31
        DATA32
        DATA33
    
'''

soup = BeautifulSoup(html)
rows = soup.findAll('tr')

section = ''
dataset = {}
for row in rows:
    if row.attrs:
        section = row.text
        dataset[section] = {}
    else:
        cells = row.findAll('td')
        for cell in cells:
            if cell['class'] in dataset[section]:
                dataset[section][ cell['class'] ].append( cell.text )
            else:
                dataset[section][ cell['class'] ] = [ cell.text ]

pprint(dataset)

Produces:

{u'HEADER 1': {u'data_point_1': [u'DATA11', u'DATA21', u'DATA31'],
               u'data_point_2': [u'DATA12', u'DATA22', u'DATA32'],
               u'data_point_3': [u'DATA12', u'DATA23', u'DATA33']},
 u'HEADER 2': {u'data_point_1': [u'DATA11', u'DATA21', u'DATA31'],
               u'data_point_2': [u'DATA12', u'DATA22', u'DATA32'],
               u'data_point_3': [u'DATA13', u'DATA23', u'DATA33']}}

EDIT ADAPTATION OF YOUR SOLUTION

Your code is neat and has only a couple of issues. You use contents in places where you shoul duse text or findAll -- I repaired that below:

soup = BeautifulSoup(html)
credits = soup.find('table')

section = ''
DATA_SET = {}

for i in credits.findAll('tr'):
    if i.get('class', '') == 'header':
        section = i.text
        DATA_SET[section] = {}
    else:
        DATA_SET[section]['data_point_1'] = i.find('td', {'class' : 'data_point_1'}).find('p').contents
        DATA_SET[section]['data_point_2'] = i.find('td', {'class' : 'data_point_2'}).find('p').contents
        DATA_SET[section]['data_point_3'] = i.find('td', {'class' : 'data_point_3'}).find('p').contents

print DATA_SET

Please note that if successive cells have the same data_point class, then successive rows will replace earlier ones. I suspect this is not an issue in your real dataset, but that is why your code would return this, abbreviated, result:

{u'HEADER 2': {'data_point_2': [u'DATA32'],
               'data_point_3': [u'DATA33'],
               'data_point_1': [u'DATA31']},
 u'HEADER 1': {'data_point_2': [u'DATA32'],
               'data_point_3': [u'DATA33'],
               'data_point_1': [u'DATA31']}}

Iterating through a table of rows with beautiful soup in python

Answers (1)

Related Questions