Pivoting a Pandas Dataframe on Categorical Variables

Question

I have a dataframe containing categorical variables:

{'SysID': {0: '00721778',
1: '00721778',
2: '00721778',
3: '00721779',
4: '00721779'},
'SoftwareComponent': {0: 'AA13912',
1: 'AA24120',
2: 'AA21612',
3: 'AA30861',
4: 'AA20635'},
'SoftwareSubcomponent': {0: None,
1: 'AK21431',
2: None,
3: 'AK22116',
4: None}}

I would like to pivot on the categorical variables by ignoring any NULL values. Zero should be the filler. The output should look like this:

{'SysID': {0: '00721778', 1: '00721779'},
'SoftwareCom-AA13912': {0: '1', 1: '0'},
'SoftwareCom-AA24120': {0: '1', 1: '0'},
'SoftwareCom-AA21612': {0: '1', 1: '0'},
'SoftwareCom-AA30861': {0: '0', 1: '1'},
'SoftwareCom-AA20635': {0: '0', 1: '1'},
'SoftwareSub-AK21431': {0: '1', 1: '0'},
'SoftwareSub-AK22116': {0: '0', 1: '1'}}

How to do this?

rahlf23 · Accepted Answer

You can use pd.crosstab() and then rename your dataframe columns prior to using pd.concat():

df1 = pd.crosstab(df['SysID'], df['SoftwareComponent'])
df1.columns = [df1.columns.name + '-' + i for i in df1.columns]
df2 = pd.crosstab(df['SysID'], df['SoftwareSubcomponent'])
df2.columns = [df2.columns.name + '-' + i for i in df2.columns]
final = pd.concat([df1, df2], axis=1)

Yields:

          SoftwareComponent-AA13912  SoftwareComponent-AA20635  \
SysID                                                            
00721778                          1                          0   
00721779                          0                          1   

          SoftwareComponent-AA21612  SoftwareComponent-AA24120  \
SysID                                                            
00721778                          1                          1   
00721779                          0                          0   

          SoftwareComponent-AA30861  SoftwareSubcomponent-AK21431  \
SysID                                                               
00721778                          0                             1   
00721779                          1                             0   

          SoftwareSubcomponent-AK22116  
SysID                                   
00721778                             0  
00721779                             1

Using to_dict(), you can return:

{'SoftwareComponent-AA13912': {'00721778': 1, '00721779': 0}, 'SoftwareComponent-AA20635': {'00721778': 0, '00721779': 1}, 'SoftwareComponent-AA21612': {'00721778': 1, '00721779': 0}, 'SoftwareComponent-AA24120': {'00721778': 1, '00721779': 0}, 'SoftwareComponent-AA30861': {'00721778': 0, '00721779': 1}, 'SoftwareSubcomponent-AK21431': {'00721778': 1, '00721779': 0}, 'SoftwareSubcomponent-AK22116': {'00721778': 0, '00721779': 1}}

Pivoting a Pandas Dataframe on Categorical Variables

Answers (2)

Output:

Related Questions