Performance issue pandas 6 mil rows

Question

need one help.

I am trying to concatenate two data frames. 1st has 58k rows, other 100. Want to concatenate in a way that each of 58k row has 100 rows from other df. So in total 5.8 mil rows. Performance is very poor, takes 1 hr to do 10 pct. Any suggestions for improvement? Here is code snippet.

def myfunc(vendors3,cust_loc):
cust_loc_vend = pd.DataFrame()
cust_loc_vend.empty
for i,row in cust_loc.iterrows():
    clear_output(wait=True)
    a= row.to_frame().T
    df= pd.concat([vendors3, a],axis=1, ignore_index=False)
    #cust_loc_vend = pd.concat([cust_loc_vend, df],axis=1, ignore_index=False)
    cust_loc_vend= cust_loc_vend.append(df)
    print('Current progress:',np.round(i/len(cust_loc)*100,2),'%')
return cust_loc_vend

For e.g. if first DF has 5 rows and second has 100 rows

DF1 (sample 2 columns)

I want a merged DF such that each row in DF 2 has All rows from DF1-

Performance issue pandas 6 mil rows

Answers (1)

Related Questions