Web scraping, python and beautifulsoup

Question

I wanted to get a paragraph from a site but ive done it this way. i get the texts of the webpage removing all html tags and i wanted to find out if its possible ta get a certain paragraph form all the text it returned.

heres my code

import requests
from bs4 import BeautifulSoup

response = requests.get("https://en.wikipedia.org/wiki/Aras_(river)")
txt = response.content

soup = BeautifulSoup(txt,'lxml')
filtered = soup.get_text()
print(filtered)

heres part of the text it printed out

>>>>Basin


    Main source
    Erzurum Province, Turkey


    River mouth
    Kura river


    Physical characteristics


    Length
    1,072 km (666 mi)


    The Aras or Araxes is a river in and along the countries of Turkey,     
    Armenia, Azerbaijan, and Iran. It drains the south side of the Lesser 
    Caucasus Mountains and then joins the Kura River which drains the north 
    side of those mountains. Its total length is 1,072 kilometres (666 mi). 
    Given its length and a basin that covers an area of 102,000 square 
    kilometres (39,000 sq mi), it is one of the largest rivers of the 
    Caucasus.



    Contents


    1 Names
    2 Description
    3 Etymology and history
    4 Iğdır Aras Valley Bird Paradise
    5 Gallery
    6 See also
    7 Footnotes

And i only want to get this paragraph

    The Aras or Araxes is a river in and along the countries of Turkey,     
    Armenia, Azerbaijan, and Iran. It drains the south side of the Lesser 
    Caucasus Mountains and then joins the Kura River which drains the north 
    side of those mountains. Its total length is 1,072 kilometres (666 mi). 
    Given its length and a basin that covers an area of 102,000 square 
    kilometres (39,000 sq mi), it is one of the largest rivers of the 
    Caucasus.

is it possible to filter out this paragraph?

宏杰李 · Accepted Answer

soup = BeautifulSoup(txt,'lxml')
filtered = soup.p.get_text() # get the first p tag.
print(filtered)

out:

The Aras or Araxes is a river in and along the countries of Turkey, Armenia, Azerbaijan, and Iran. It drains the south side of the Lesser Caucasus Mountains and then joins the Kura River which drains the north side of those mountains. Its total length is 1,072 kilometres (666 mi). Given its length and a basin that covers an area of 102,000 square kilometres (39,000 sq mi), it is one of the largest rivers of the Caucasus.

Web scraping, python and beautifulsoup

Answers (2)

Related Questions