XPath - extracting text between two nodes

Question

I'm encountering a problem with my XPath query. I have to parse a div which is divided to unknown number of "sections". Each of these is separated by h5 with a section name. The list of possible section titles is known and each of them can occur only once. Additionally, each section can contain some br tags. So, let's say I want to extract the text under "SecondHeader".

HTML


 FirstHeader
  text1
 SecondHeader
  text2a

  text2b
 ThirdHeader
  text3a

  text3b

  text3c

 FourthHeader
  text4

Expected result (for SecondSection)

['text2a', 'text2b']

Query #1

//text()[following-sibling::h5/text()='ThirdHeader']

Result #1

['text1', 'text2a', 'text2b']

It's obviously bit too much, so I've decided to restrict the result to the content between selected header and the header before.

Query #2

//text()[following-sibling::h5/text()='ThirdHeader' and preceding-sibling::h5/text()='SecondHeader']

Result #2

['text2a', 'text2b']

Yielded results meet the expectations. However, this can't be used - I don't know whether SecondHeader/ThirdHeader will exist in parsed page or not. It is needed to use only one section title in a query.

Query #3

//text()[following-sibling::h5/text()='ThirdHeader' and not[preceding-sibling::h5/text()='ThirdHeader']]

Result #3

[]

Could you please tell me what am I doing wrong? I've tested it in Google Chrome.

Daniel Haley · Accepted Answer

You should be able to just test the first preceding sibling h5...

//text()[preceding-sibling::h5[1][normalize-space()='SecondHeader']]

XPath - extracting text between two nodes

Answers (2)

Related Questions