Home »
MCQs
Web Scraping with Python MCQs (Multiple-Choice Questions)
Web scraping with Python is the process of programmatically retrieving and extracting information from web pages. Python provides several libraries and tools for web scraping, including Requests, urllib, Beautiful Soup, Selenium, and other specialized libraries. These tools can be used to send HTTP requests, parse HTML documents, locate elements, extract text and links, and automate browser interactions.
Web Scraping with Python MCQs
These Web Scraping with Python MCQs cover HTTP requests, Requests, urllib, Beautiful Soup, HTML parsing, CSS selectors, XPath, headers, sessions, cookies, status codes, robots.txt, pagination, dynamic content, Selenium, and practical web scraping techniques.
List of Web Scraping with Python MCQs
Practice these multiple-choice questions to test your knowledge of web scraping concepts and Python libraries used for extracting data from websites.
1. What is web scraping?
- Creating a website using Python
- Programmatically extracting data from web pages
- Designing a database schema
- Compiling HTML documents
Answer: B) Programmatically extracting data from web pages
Explanation:
Web scraping involves retrieving information from web resources and extracting useful data programmatically rather than manually copying it.
2. Which Python library is commonly used to send HTTP requests for web scraping?
- Requests
- NumPy
- Matplotlib
- tkinter
Answer: A) Requests
Explanation:
Requests is a popular Python HTTP client that provides a convenient API for sending HTTP requests and handling responses.
3. Which Requests function is commonly used to perform an HTTP GET request?
requests.fetch()
requests.get()
requests.retrieve()
requests.open()
Answer: B) requests.get()
Explanation:
requests.get() sends an HTTP GET request to the specified URL and returns a Requests Response object.
4. Which HTTP method is primarily used to retrieve a resource from a web server?
- GET
- POST
- DELETE
- PATCH
Answer: A) GET
Explanation:
HTTP GET is generally used to request and retrieve a resource from a server. It is commonly used when downloading HTML pages for scraping.
5. Which Requests attribute contains the response body as decoded text?
response.html
response.text
response.body_text
response.document
Answer: B) response.text
Explanation:
The text attribute provides the response content as a decoded string. Requests also provides content for the response body as bytes.
6. Which Requests attribute provides the response body as bytes?
response.bytes
response.raw_text
response.content
response.data
Answer: C) response.content
Explanation:
response.content returns the response body as bytes, which is useful when working with binary resources such as images or when decoding content manually.
7. Which Requests attribute contains the HTTP status code?
response.code
response.status
response.status_code
response.http_code
Answer: C) response.status_code
Explanation:
The status_code attribute contains the HTTP response status code, such as 200, 404, or 500.
8. What does HTTP status code 200 generally indicate?
- The resource was permanently moved
- The request was successful
- The server encountered an error
- The resource was not found
Answer: B) The request was successful
Explanation:
HTTP 200 is a successful response status indicating that the request was successfully processed.
9. What does HTTP status code 404 normally indicate?
- Unauthorized access
- Successful response
- Resource not found
- Too many requests
Answer: C) Resource not found
Explanation:
An HTTP 404 response indicates that the requested resource could not be found on the server.
10. Which Requests method raises an HTTPError for unsuccessful HTTP status codes?
response.check_status()
response.raise_for_status()
response.validate()
response.raise_error()
Answer: B) response.raise_for_status()
Explanation:
raise_for_status() raises an HTTPError when the response contains an unsuccessful HTTP status code such as a 4xx or 5xx response.
11. Why should a scraper generally specify a timeout for Requests calls?
- To increase HTML size
- To prevent the request from waiting indefinitely
- To disable HTTPS
- To force every response to return status 200
Answer: B) To prevent the request from waiting indefinitely
Explanation:
Requests does not time out automatically unless a timeout is specified. Setting a timeout helps prevent a scraper from waiting indefinitely for an unresponsive server.
12. Which parameter is used to specify a timeout in requests.get()?
delay
timeout
wait
max_time
Answer: B) timeout
Explanation:
The timeout parameter controls how long Requests waits for network activity before raising a timeout exception.
13. Which parameter is used to send custom HTTP headers with a Requests call?
headers
http_headers
request_headers
metadata
Answer: A) headers
Explanation:
Custom HTTP headers can be supplied using the headers parameter, for example to specify a User-Agent.
14. What is the purpose of the HTTP User-Agent header in a web request?
- It identifies the requested database
- It identifies the client software making the request
- It encrypts the response
- It specifies the HTML parser
Answer: B) It identifies the client software making the request
Explanation:
The User-Agent header identifies the client software making an HTTP request, such as a browser, crawler, or custom application.
15. Which Python library provides urllib.request for opening URLs?
- urllib
- urltools
- weburllib
- httpurl
Answer: A) urllib
Explanation:
The Python standard library's urllib package contains modules such as urllib.request, urllib.parse, and urllib.robotparser.
16. Which function from urllib.request can retrieve a URL?
urlopen()
fetchurl()
geturl()
download_url()
Answer: A) urlopen()
Explanation:
urllib.request.urlopen() opens a URL and returns a response object that can be used to read the resource.
17. Which library is commonly used to parse HTML documents in Python?
- Beautiful Soup
- NumPy
- PyGame
- Flask
Answer: A) Beautiful Soup
Explanation:
Beautiful Soup is an HTML and XML parsing library that provides convenient APIs for navigating and searching parsed documents.
18. Which import correctly imports Beautiful Soup's main parser class?
from bs4 import BeautifulSoup
from beautifulsoup import Parser
import BeautifulSoup from bs4
from soup import BeautifulSoup
Answer: A) from bs4 import BeautifulSoup
Explanation:
Beautiful Soup 4 is installed as the bs4 package, and BeautifulSoup is imported from it.
19. Which statement creates a Beautiful Soup object from HTML stored in html?
BeautifulSoup(html, "html.parser")
BeautifulSoup.parse(html)
Soup(html, "html")
HTMLParser.parse(html)
Answer: A) BeautifulSoup(html, "html.parser")
Explanation:
The BeautifulSoup constructor accepts markup and a parser specification. Python's built-in html.parser can be used for basic HTML parsing.
20. Which Beautiful Soup method finds all matching elements in a document?
find_all()
find_every()
search_all()
get_all()
Answer: A) find_all()
Explanation:
find_all() searches the document tree and returns all elements matching the supplied filters.
21. Which Beautiful Soup method returns the first matching element?
find()
first()
find_one()
search()
Answer: A) find()
Explanation:
find() returns the first matching element, or None when no matching element is found.
22. Which Beautiful Soup expression finds all anchor elements?
soup.find_all("a")
soup.get_links("a")
soup.links()
soup.find_links("a")
Answer: A) soup.find_all("a")
Explanation:
The string "a" specifies the HTML anchor tag, and find_all() returns all matching tags.
23. Which Beautiful Soup method extracts all text contained in a document or tag?
get_text()
extract_text()
text_all()
plain_text()
Answer: A) get_text()
Explanation:
get_text() returns the text contained within a Beautiful Soup document or tag as a string.
24. Which expression retrieves the value of an HTML anchor's href attribute?
link.href
link.get("href")
link.attribute("href")
link.url()
Answer: B) link.get("href")
Explanation:
Beautiful Soup Tag objects provide get() for safely retrieving an attribute value. If the attribute does not exist, it can return None.
25. Which Beautiful Soup method supports CSS selectors?
select()
css_find()
query_css()
find_css_all()
Answer: A) select()
Explanation:
Beautiful Soup's select() method uses CSS selectors to locate matching elements in the parsed document.
26. What does the CSS selector ".product" generally select?
- Elements with class "product"
- The element with ID "product"
- All product tags
- Elements named "product"
Answer: A) Elements with class "product"
Explanation:
In CSS selectors, a period followed by a name represents a class selector. Therefore, .product selects elements whose class includes product.
27. What does the CSS selector "#price" select?
- Elements with class "price"
- The element with ID "price"
- All elements named price
- All price attributes
Answer: B) The element with ID "price"
Explanation:
In CSS selectors, the # prefix represents an ID selector.
28. Which Beautiful Soup parser uses Python's built-in HTML parser?
html.parser
python.html
builtin.html
native.parser
Answer: A) html.parser
Explanation:
html.parser is Python's built-in HTML parser and can be passed to Beautiful Soup as the parser specification.
29. What is the primary purpose of XPath in web scraping?
- To encrypt HTTP traffic
- To locate nodes in XML or HTML documents
- To compress HTML pages
- To create HTTP headers
Answer: B) To locate nodes in XML or HTML documents
Explanation:
XPath is a query language for navigating and selecting nodes in XML-like document structures. It is commonly used with browser automation and HTML parsing tools that support XPath.
30. Why might a simple Requests request fail to retrieve data visible in a browser?
- Requests cannot download HTML
- The data may be generated dynamically by JavaScript after the initial HTML response
- HTML cannot be sent over HTTP
- Beautiful Soup automatically blocks Requests
Answer: B) The data may be generated dynamically by JavaScript after the initial HTML response
Explanation:
Requests retrieves the HTTP response but does not execute JavaScript like a browser. If important content is generated after page load through JavaScript, a browser automation tool or direct underlying API may be needed.
31. Which tool is commonly used with Python to automate a real web browser?
- Selenium
- NumPy
- Beautiful Soup
- SQLite
Answer: A) Selenium
Explanation:
Selenium provides browser automation capabilities and can interact with web pages in a real browser environment, making it useful for pages that require JavaScript execution or browser interaction.
32. What is a headless browser?
- A browser that runs without a visible graphical user interface
- A browser without HTTP support
- A browser that cannot execute JavaScript
- A browser that only supports HTML files
Answer: A) A browser that runs without a visible graphical user interface
Explanation:
A headless browser runs browser functionality without displaying a normal graphical browser window. It can be useful for automated scraping and testing.
33. Which approach is generally preferable when a website exposes the required data through a documented API?
- Use the API instead of scraping rendered HTML when appropriate
- Disable the website's security mechanisms
- Send thousands of requests simultaneously
- Ignore the API and scrape every page
Answer: A) Use the API instead of scraping rendered HTML when appropriate
Explanation:
A documented API usually provides structured data and a defined interface, which can be more reliable and appropriate than parsing rendered HTML.
34. What is the purpose of a web page's robots.txt file?
- To define HTML styles
- To communicate crawler access rules
- To store JavaScript variables
- To encrypt page content
Answer: B) To communicate crawler access rules
Explanation:
A robots.txt file provides rules that automated clients can use to determine which URLs a site's robots policy permits or disallows for specified user agents.
35. Which Python class can be used to evaluate robots.txt rules?
urllib.robotparser.RobotFileParser
urllib.robots.RobotParser
urllib.request.RobotParser
urllib.crawler.Robots
Answer: A) urllib.robotparser.RobotFileParser
Explanation:
Python's urllib.robotparser module provides RobotFileParser for reading and evaluating robots.txt rules.
36. Which method of RobotFileParser checks whether a user agent can fetch a URL?
can_fetch()
is_allowed()
allowed_url()
check_access()
Answer: A) can_fetch()
Explanation:
The can_fetch() method returns whether a specified user agent is allowed to fetch a URL according to the parsed robots.txt rules.
37. Why is adding a delay between scraping requests sometimes useful?
- It makes HTML invalid
- It can reduce request frequency and server load
- It guarantees that the server will not block the scraper
- It converts HTTP to HTTPS
Answer: B) It can reduce request frequency and server load
Explanation:
A controlled request rate can reduce load on the target server and make a scraper less aggressive. A delay does not guarantee that a scraper will never be blocked.
38. Which Requests object can preserve cookies and other session-related state across requests?
requests.Session()
requests.CookieJar()
requests.State()
requests.Persistent()
Answer: A) requests.Session()
Explanation:
A Requests Session object persists certain parameters and cookies across requests, making it useful when a scraper needs to maintain session state.
39. Which parameter is commonly used with Requests to send query-string parameters?
params
query
url_params
search
Answer: A) params
Explanation:
The params parameter can be used to pass query-string parameters to a Requests GET request.
40. What is pagination in web scraping?
- Converting HTML into CSS
- Processing data distributed across multiple pages
- Compressing downloaded pages
- Changing a server's database
Answer: B) Processing data distributed across multiple pages
Explanation:
Pagination is a common website pattern in which a collection of results is split across multiple pages. A scraper may need to follow page links or construct subsequent page URLs.
41. Which HTML element commonly contains a link that can be followed to the next page of paginated results?
<meta>
<a>
<title>
<style>
Answer: B) <a>
Explanation:
An anchor element commonly contains the hyperlink used to navigate to another page, including a next-page URL.
42. Why may a scraper need to convert a relative link into an absolute URL?
- Relative URLs cannot be stored as strings
- An absolute URL is required to make a complete request independently of the current page path
- Relative URLs are always invalid HTML
- Requests only supports IP addresses
Answer: B) An absolute URL is required to make a complete request independently of the current page path
Explanation:
HTML frequently contains relative links such as /products/1. A scraper can resolve these against the source page URL before making another request.
43. Which Python module provides URL parsing and joining functionality?
urllib.parse
urllib.html
urllib.urls
url.parser
Answer: A) urllib.parse
Explanation:
urllib.parse provides functions for parsing, constructing, and manipulating URLs and related components.
44. Which HTTP status code commonly indicates that the client has sent too many requests in a given period?
- 301
- 403
- 429
- 503
Answer: C) 429
Explanation:
HTTP 429 means Too Many Requests. A scraper encountering this response should respect the target site's policies and implement appropriate rate limiting or retry behavior when permitted.
45. Which HTTP status code indicates that the server understood the request but refuses to authorize it?
- 200
- 301
- 403
- 404
Answer: C) 403
Explanation:
HTTP 403 Forbidden indicates that the server understood the request but refuses to authorize access to the requested resource.
46. What is the main limitation of using a regular expression as the primary HTML parser?
- Regular expressions cannot process strings
- HTML has nested and irregular document structures that dedicated parsers handle more appropriately
- Regular expressions cannot search text
- Regular expressions only work with JSON
Answer: B) HTML has nested and irregular document structures that dedicated parsers handle more appropriately
Explanation:
HTML is a structured markup language with nested elements and potentially malformed markup. Dedicated parsers such as Beautiful Soup are designed to navigate and extract information from HTML documents.
47. Which exception from Requests is associated with a network-level connection failure?
requests.exceptions.ConnectionError
requests.exceptions.NetworkFailure
requests.exceptions.ServerError
requests.exceptions.SocketFailure
Answer: A) requests.exceptions.ConnectionError
Explanation:
Requests raises ConnectionError when a network problem prevents the connection from being successfully established or maintained.
48. What does response.json() attempt to do in Requests?
- Convert HTML into XML
- Decode the response body as JSON
- Extract all hyperlinks
- Convert JSON into HTML
Answer: B) Decode the response body as JSON
Explanation:
response.json() attempts to decode the response content as JSON. A successful JSON decode does not by itself mean that the HTTP request succeeded, so the status code should also be checked.
49. Consider the following code. What does it print if the page contains three <a> elements?
import requests
from bs4 import BeautifulSoup
response = requests.get("https://example.com", timeout=10)
soup = BeautifulSoup(response.text, "html.parser")
links = soup.find_all("a")
print(len(links))
- The number of links found by Beautiful Soup
- The HTTP status code
- The length of the downloaded HTML in bytes
- The URL of the page
Answer: A) The number of links found by Beautiful Soup
Explanation:
find_all("a") returns a collection of matching anchor elements. Therefore, len(links) gives the number of anchor elements found in the parsed document.
50. Consider the following code. What will title contain if the HTML contains <title>Python Scraping</title>?
from bs4 import BeautifulSoup
html = "Python Scraping"
soup = BeautifulSoup(html, "html.parser")
title = soup.title.get_text(strip=True)
title
<title>
Python Scraping
None
Answer: C) Python Scraping
Explanation:
soup.title accesses the title tag, and get_text(strip=True) extracts its text while removing surrounding whitespace.
Advertisement
Advertisement