Tripadvisor 的 Scrapy Spider 抓取了 0 页(以 0 页/分钟的速度) [英] Scrapy spider for Tripadvisor Crawled 0 pages (at 0 pages/min)

查看:52
本文介绍了Tripadvisor 的 Scrapy Spider 抓取了 0 页(以 0 页/分钟的速度)的处理方法,对大家解决问题具有一定的参考价值,需要的朋友们下面随着小编来一起学习吧!

问题描述

我正在尝试收集蓬塔卡纳所有酒店的评论.代码似乎可以运行,但是当我调用 crawl 时,它实际上并没有爬取任何站点.这是我的文件结构、我调用的内容以及运行时发生的情况.

文件夹结构:

├── scrapy.cfg└── tripadvisor_reviews├── __init__.py├── __pycache__│ ├── __init__.cpython-37.pyc│ ├── items.cpython-37.pyc│ └── settings.cpython-37.pyc├──物品.py├── 中间件.py├── 管道.py├── settings.py└── 蜘蛛├── __init__.py├── __pycache__│ ├── __init__.cpython-37.pyc│ └── tripadvisorSpider.cpython-37.pyc└── tripadvisorSpider.py

tripadvisorSpider.py

导入scrapy从 tripadvisor_reviews.items 导入 TripadvisorReviewsItem类tripadvisorSpider(scrapy.Spider):名称 = "tripadvisorspider"allowed_domains = ["www.tripadvisor.com"]def start_requests(self):网址 = ['https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html']对于网址中的网址:yield scrapy.Request(url=url, callback=self.parse)定义解析(自我,响应):对于 response.xpath 中的 href('//div[@class="listing_title"]/a/@href'):url = response.urljoin(href.extract())产生scrapy.Request(url, callback=self.parse_hotel)next_page = response.xpath('//div[@class="nav next taLnk ui_button primary"]/@href').extract_first()如果下一页:url = response.urljoin(next_page)产生scrapy.Request(url,self.parse)def parse_hotel(self, response):对于 href 在 response.xpath('//div[@class="hotels-review-list-parts-ReviewTitle__reviewTitleText"]/a/@href'):url = response.urljoin(href.extract())产生scrapy.Request(url,回调=self.parse_review)next_page = response.xpath('//div[@class="ui_button nav next primary"]/@href').extract_first()如果下一页:url = response.urljoin(next_page)产生scrapy.Request(网址,self.parse_hotel)def parse_review(self, response):item = TripadvisorReviewsItem()item['title'] = response.xpath('//div[@class="hotels-review-list-parts-ReviewTitle__reviewTitleText"]/text()').extract()item['content'] = response.xpath('//q[@class="hotels-review-list-parts-ExpandableReview__reviewText"]/text()').extract()# item['stars'] = response.xpath(# '//span[@class="rate sprite-rating_s rating_s"]/img/@alt').extract()[0]打印(项目)产量项目

items.py

导入scrapy类 TripadvisorReviewsItem(scrapy.Item):# 在此处为您的项目定义字段,例如:标题 = scrapy.Field()内容=scrapy.Field()# 星星 = scrapy.Field()

我在终端中使用以下命令运行它:

scrapy 爬取tripadvisorspider -o items.json

这是我的终端输出

2019-05-14 12:32:12 [scrapy.utils.log] 信息:Scrapy 1.5.2 开始(机器人:tripadvisor_reviews)2019-05-14 12:32:12 [scrapy.utils.log] 信息:版本:lxml 4.2.5.0、libxml2 2.9.8、cssselect 1.0.3、parsel 1.5.1、w3lib 1.20.0、Twisted 19.2.2, Python 3.7.1 (default, Dec 14 2018, 13:28:58) - [Clang 4.0.1 (tags/RELEASE_401/final)], pyOpenSSL 18.0.0 (OpenSSL 1.1.1b 26 Feb 2019), 密码学 2.4.2、平台Darwin-18.5.0-x86_64-i386-64bit2019-05-14 12:32:12 [scrapy.crawler] 信息:覆盖设置:{'BOT_NAME':'tripadvisor_reviews','FEED_FORMAT':'csv','FEED_URI':'items.csv','NEWSPIDER_MODULE':'tripadvisor_reviews.spiders','ROBOTSTXT_OBEY':真,'SPIDER_MODULES':['tripadvisor_reviews.spiders']}2019-05-14 12:32:12 [scrapy.extensions.telnet] 信息:Telnet 密码:aae78556d6b8c59b2019-05-14 12:32:12 [scrapy.middleware] 信息:启用扩展:['scrapy.extensions.corestats.CoreStats','scrapy.extensions.telnet.TelnetConsole','scrapy.extensions.memusage.MemoryUsage','scrapy.extensions.feedexport.FeedExporter','scrapy.extensions.logstats.LogStats']2019-05-14 12:32:12 [scrapy.middleware] 信息:启用下载器中间件:['scrapy.downloadermiddlewares.robotstxt.RobotsTxtMiddleware','scrapy.downloadermiddlewares.httpauth.HttpAuthMiddleware','scrapy.downloadermiddlewares.downloadtimeout.DownloadTimeoutMiddleware','scrapy.downloadermiddlewares.defaultheaders.DefaultHeadersMiddleware','scrapy.downloadermiddlewares.useragent.UserAgentMiddleware','scrapy.downloadermiddlewares.retry.RetryMiddleware','scrapy.downloadermiddlewares.redirect.MetaRefreshMiddleware','scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware','scrapy.downloadermiddlewares.redirect.RedirectMiddleware','scrapy.downloadermiddlewares.cookies.CookiesMiddleware','scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware','scrapy.downloadermiddlewares.stats.DownloaderStats']2019-05-14 12:32:12 [scrapy.middleware] 信息:启用蜘蛛中间件:['scrapy.spidermiddlewares.httperror.HttpErrorMiddleware','scrapy.spidermiddlewares.offsite.OffsiteMiddleware','scrapy.spidermiddlewares.referer.RefererMiddleware','scrapy.spidermiddlewares.urllength.UrlLengthMiddleware','scrapy.spidermiddlewares.depth.DepthMiddleware']2019-05-14 12:32:12 [scrapy.middleware] 信息:启用项目管道:[]2019-05-14 12:32:12 [scrapy.core.engine] 信息:Spider 打开2019-05-14 12:32:12 [scrapy.extensions.logstats] 信息:抓取 0 页(以 0 页/分钟),抓取 0 个项目(以 0 个项目/分钟)2019-05-14 12:32:12 [scrapy.extensions.telnet] 调试:Telnet 控制台监听 127.0.0.1:60242019-05-14 12:32:16 [scrapy.core.engine] 调试:爬行(200)<GET https://www.tripadvisor.com/robots.txt>(参考:无)2019-05-14 12:32:17 [scrapy.core.engine] 调试:爬行(200)<GET https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html>(参考:无)2019-05-14 12:32:17 [scrapy.core.engine] DEBUG:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g3176298-d15025732-Reviews-Impressive_Resort_ana_Punta_PuntaC(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:18 [scrapy.core.engine] 调试:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g3176298-d313884-Reviews-Punta_Cana_Princess_All_Spaint_Proc_Al_Spaña_Proc(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:18 [scrapy.core.engine] 调试:爬行 (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d4451011-Reviews-The_Westin_Puntacana_Resort_AltaC(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:18 [scrapy.core.engine] 调试:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g147293-d10175054-Reviews-Secrets_Cap_Cana_Resort_Vince_Alccia_Alce_Alce_CanaP(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:18 [scrapy.core.engine] DEBUG:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g3176298-d7307251-Reviews-The_Level_at_Melia_Caribero_VintaReviews-The_Level_at_Melia_Caribe_Vinta_Caribe_Alce_Pro_VinciaP(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:18 [scrapy.core.engine] 调试:爬行 (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d292158-Reviews-Grand_Palladium_Punta_Cana_Punta_Cana_Punta_Cana_PuntaGrand_Palladium_Punta_Cana_Puntac(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:18 [scrapy.core.engine] DEBUG:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g147293-d1604057-Reviews-Secrets_Royal_Beach_Punta_Reviews-Secrets_Royal_Beach_Punta_Punta_Punta_Crawled(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:18 [scrapy.core.engine] 调试:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g147293-d649099-Reviews-Zoetry_Agua_Punta_Cana_Dominta_Cana_Punta_Punta_Punta_Cana_Punta<GET https://www.tripadvisor.com/Hotel_Review-g147293-d649099-(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:18 [scrapy.core.engine] DEBUG:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g3176298-d150842-Reviews-Iberostar_Dominicana_Hotel-Alvinagramini<GET https://www.tripadvisor.com/Hotel_Review-g3176298-d150842-Reviews-Iberostar_Dominicana_Hotel-Bavaro_p(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:19 [scrapy.core.engine] 调试:爬行 (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d15515013-Reviews-Grand_Memories_Punta_Cana_Review_Punta_Cana-Reviews_Punta_Cana_AlcenaPunta_Cana-Reviews(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:19 [scrapy.core.engine] 调试:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g3176298-d150841-Reviews-Iberostar_Selection_Bavaro_AlvinCiaPavaro_AlvinCiaPavaro_AlvinCaminiReviews(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:19 [scrapy.core.engine] DEBUG:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g3176298-d1233228-Reviews-Iberostar_Grand_Bavaro_Pavaro_Altavaro_Pavaro_Altavaro<GET https://www.tripadvisor.com/Hotel_Review-g3176298-d1233228-Reviews(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:19 [scrapy.core.engine] 调试:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g147293-d149397-Reviews-Bavaro_Princess_Resort_Spa_Casino_Resort_Spa_Casino-htmltciaPara(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:19 [scrapy.core.engine] 调试:爬行 (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d584407-Reviews-Ocean_Blue_Sand-Bavaro_and-Bavaro_and-Bavaro_PuntaReview_Alvint(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:19 [scrapy.core.engine] 调试:爬行 (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d259337-Reviews-Grand_Palladium_Bavaro_Suites_Reviews-Grand_Palladium_Bavaro_Suites_Reviews-Grand_Palladium_Bavaro_Suites_Reviews(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:19 [scrapy.core.engine] DEBUG:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g147293-d150854-Reviews-Hotel_Riu_Palace_Macao-Palace_Macao-Punta_Alce_Alce_Macao-Punta_Review_Alt;(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:20 [scrapy.core.engine] DEBUG:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g147293-d11701188-Reviews-BlueBay_Grand_Punta_Cana_Alce_Altcia_Punta_Cana-Punta_Cana-Punta_Cana-Punt(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:20 [scrapy.core.engine] 调试:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g147293-d14838260-Reviews-Melia_Punta_Cana_Beach_Reviews-Melia_Punta_Cana_Beach_Resort_Html_Beach_Reviews(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:20 [scrapy.core.engine] 调试:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g147293-d1595124-Reviews-Luxury_Bahia_Principe_Principe_Principe_Principe_Principe_Principe_Principe_Principe_HtmlLauntameralc(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:20 [scrapy.core.engine] DEBUG:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g147293-d508162-Reviews-Dreams_Punta_Cana_Resort_Vinta-Cana_Resort_spa-anaPunta-Crawled<GET https://www.tripadvisor.com/Hotel_Review-g147293-d508162-(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:20 [scrapy.core.engine] DEBUG:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g3176298-d1076311-Reviews-Hard_Rock_Hotel_Casino_Pavarot_Crawl_Punta_Alvince_Casino_Punta_Crawl(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:20 [scrapy.core.engine] 调试:爬行 (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d10595200-Reviews-Grand_Bahia_Principe_Alvincia_Principe_HtmlPrincipe_Aquamart(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:20 [scrapy.core.engine] 调试:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g147293-d611114-Reviews-Hotel_Riu_Palace_Punta_Cana_Review_Punta_CanaPunta_CanaPunta_CanaPunta<GET https://www.tripadvisor.com/Hotel_Review-g147293-d6111114(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:20 [scrapy.core.engine] 调试:爬行 (200) <GET https://www.tripadvisor.com/Hotel_Review-g3200043-d8709413-Reviews-Excellence_El_Carmen-Uvero_Altagran_PublicReviews-Excellence_El_Carmen-Uvero_Lavintac(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:21 [scrapy.core.engine] 调试:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g147293-d1199681-Reviews-Luxury_Bahia_Principe_Principe_Altaar<GET(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:21 [scrapy.core.engine] 调试:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g3176298-d1889895-Reviews-Karibo_Punta_Cana_Pavaro_PavaroCana_Pavaro<GET https://www.tripadvisor.com/Hotel_Review-g3176298-d1889895-Reviews-Karibo_Punta_Cana_Pavaro_htmlReviews(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:21 [scrapy.core.engine] 调试:爬行 (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d6454132-Reviews-Premium_Level_at_Barcelo_Pavaroc_Alcia_Pavaro<获取 https://www.tripadvisor.com/Hotel_Review-g3176298-d6454132-Reviews-Premium_Level_at_Barcelo_Pavarod(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:21 [scrapy.core.engine] 调试:爬行 (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d15080584-Reviews-Impressive_Premium_Resortvin_Resort_Alce_Alce_Resort_Resort_Resort_Resort_Resort_Resort_Resort_Resort_Reviews(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:21 [scrapy.core.engine] 调试:爬行 (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d2687221-Reviews-NH_Punta_Cana-Punta_Provinta_Cana_Reviews-Nh_Punta_Cana-Punta_Alce_Alce_Alce_Alce_Alce_Lamma_La(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:21 [scrapy.core.engine] 调试:爬行(200)<GET https://www.tripadvisor.com/Hotel_Review-g3176298-d579774-Reviews-Iberostar_Punta_Punta_Panta_Pana-Bavaro_Reviews-Iberostar_Punta_Punta_Pana_Pavaro_Pavaro<get https://www.tripadvisor.com/Hotel_Review-g3176298-d579774-Reviews(参考:https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)2019-05-14 12:32:21 [scrapy.core.engine] 信息:关闭蜘蛛(已完成)2019-05-14 12:32:21 [scrapy.statscollectors] 信息:倾销 Scrapy 统计信息:{'下载器/请求字节':46023,'下载者/请求计数':32,'下载器/request_method_count/GET':32,下载器/响应字节":5599418,'下载者/响应计数':32,'下载器/response_status_count/200':32,'finish_reason': '完成','finish_time': datetime.datetime(2019, 5, 14, 19, 32, 21, 637712),'log_count/DEBUG': 33,'log_count/INFO': 8,'memusage/max':51412992,'memusage/启动':51412992,'request_depth_max': 1,'response_received_count':32,'调度程序/出队':31,调度程序/出队/内存":31,'调度程序/排队':31,调度程序/排队/内存":31,'start_time': datetime.datetime(2019, 5, 14, 19, 32, 12, 996979)}2019-05-14 12:32:21 [scrapy.core.engine] 信息:Spider 关闭(已完成)

解决方案

此选择器不起作用:

response.xpath('//div[@class="hotels-review-list-parts-ReviewTitle__reviewTitleText"]/a/@href')

网站上的元素是 而不是

,而且类名似乎是错误的.也许这个站点在类名后附加了一些随机数据,如下所示

<a href="/ShowUserReviews-g3176298-d259337-r673990694-Grand_Palladium_Bavaro_Suites_Resort_Spa-Bavaro_Punta_Cana_La_Altagracia_Provinc.html" class="hotels">Tirespanr3" class="hotels">Tirespanr3;span>Ótima体验!Resort amplo, com diversas opções de entretenimento!!!</span></span></a>

您可以尝试只匹配字符串的一部分,例如:

//a[contains(@class, "hotels-review-list-parts-ReviewTitle__reviewTitleText")]

I'm trying to scrape reviews for all hotels in Punta Cana. The code seems to run but when I call crawl, it doesn't actually crawl any of the sites. Here are my file structures, what I called, and what happened when I ran it.

folder structure:

├── scrapy.cfg
└── tripadvisor_reviews
    ├── __init__.py
    ├── __pycache__
    │   ├── __init__.cpython-37.pyc
    │   ├── items.cpython-37.pyc
    │   └── settings.cpython-37.pyc
    ├── items.py
    ├── middlewares.py
    ├── pipelines.py
    ├── settings.py
    └── spiders
        ├── __init__.py
        ├── __pycache__
        │   ├── __init__.cpython-37.pyc
        │   └── tripadvisorSpider.cpython-37.pyc
        └── tripadvisorSpider.py

tripadvisorSpider.py

import scrapy
from tripadvisor_reviews.items import TripadvisorReviewsItem


class tripadvisorSpider(scrapy.Spider):
    name = "tripadvisorspider"
    allowed_domains = ["www.tripadvisor.com"]

    def start_requests(self):

        urls = [
            'https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html'
        ]

        for url in urls:
            yield scrapy.Request(url=url, callback=self.parse)

    def parse(self, response):
        for href in response.xpath('//div[@class="listing_title"]/a/@href'):
            url = response.urljoin(href.extract())
            yield scrapy.Request(url, callback=self.parse_hotel)

        next_page = response.xpath(
            '//div[@class="nav next taLnk ui_button primary"]/@href').extract_first()
        if next_page:
            url = response.urljoin(next_page)
            yield scrapy.Request(url, self.parse)

    def parse_hotel(self, response):
        for href in response.xpath('//div[@class="hotels-review-list-parts-ReviewTitle__reviewTitleText"]/a/@href'):
            url = response.urljoin(href.extract())
            yield scrapy.Request(url, callback=self.parse_review)

        next_page = response.xpath(
            '//div[@class="ui_button nav next primary "]/@href').extract_first()
        if next_page:
            url = response.urljoin(next_page)
            yield scrapy.Request(url, self.parse_hotel)

    def parse_review(self, response):
        item = TripadvisorReviewsItem()
        item['title'] = response.xpath(
            '//div[@class="hotels-review-list-parts-ReviewTitle__reviewTitleText"]/text()').extract()
        item['content'] = response.xpath(
            '//q[@class="hotels-review-list-parts-ExpandableReview__reviewText"]/text()').extract()
        # item['stars'] = response.xpath(
        #     '//span[@class="rate sprite-rating_s rating_s"]/img/@alt').extract()[0]
        print(item)
        yield item

items.py

import scrapy


class TripadvisorReviewsItem(scrapy.Item):
    # define the fields for your item here like:
    title = scrapy.Field()
    content = scrapy.Field()
    # stars = scrapy.Field()

I ran it using the following command in terminal:

scrapy crawl tripadvisorspider -o items.json

This is my terminal output

2019-05-14 12:32:12 [scrapy.utils.log] INFO: Scrapy 1.5.2 started (bot: tripadvisor_reviews)
2019-05-14 12:32:12 [scrapy.utils.log] INFO: Versions: lxml 4.2.5.0, libxml2 2.9.8, cssselect 1.0.3, parsel 1.5.1, w3lib 1.20.0, Twisted 19.2.0, Python 3.7.1 (default, Dec 14 2018, 13:28:58) - [Clang 4.0.1 (tags/RELEASE_401/final)], pyOpenSSL 18.0.0 (OpenSSL 1.1.1b  26 Feb 2019), cryptography 2.4.2, Platform Darwin-18.5.0-x86_64-i386-64bit
2019-05-14 12:32:12 [scrapy.crawler] INFO: Overridden settings: {'BOT_NAME': 'tripadvisor_reviews', 'FEED_FORMAT': 'csv', 'FEED_URI': 'items.csv', 'NEWSPIDER_MODULE': 'tripadvisor_reviews.spiders', 'ROBOTSTXT_OBEY': True, 'SPIDER_MODULES': ['tripadvisor_reviews.spiders']}
2019-05-14 12:32:12 [scrapy.extensions.telnet] INFO: Telnet Password: aae78556d6b8c59b
2019-05-14 12:32:12 [scrapy.middleware] INFO: Enabled extensions:
['scrapy.extensions.corestats.CoreStats',
 'scrapy.extensions.telnet.TelnetConsole',
 'scrapy.extensions.memusage.MemoryUsage',
 'scrapy.extensions.feedexport.FeedExporter',
 'scrapy.extensions.logstats.LogStats']
2019-05-14 12:32:12 [scrapy.middleware] INFO: Enabled downloader middlewares:
['scrapy.downloadermiddlewares.robotstxt.RobotsTxtMiddleware',
 'scrapy.downloadermiddlewares.httpauth.HttpAuthMiddleware',
 'scrapy.downloadermiddlewares.downloadtimeout.DownloadTimeoutMiddleware',
 'scrapy.downloadermiddlewares.defaultheaders.DefaultHeadersMiddleware',
 'scrapy.downloadermiddlewares.useragent.UserAgentMiddleware',
 'scrapy.downloadermiddlewares.retry.RetryMiddleware',
 'scrapy.downloadermiddlewares.redirect.MetaRefreshMiddleware',
 'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware',
 'scrapy.downloadermiddlewares.redirect.RedirectMiddleware',
 'scrapy.downloadermiddlewares.cookies.CookiesMiddleware',
 'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware',
 'scrapy.downloadermiddlewares.stats.DownloaderStats']
2019-05-14 12:32:12 [scrapy.middleware] INFO: Enabled spider middlewares:
['scrapy.spidermiddlewares.httperror.HttpErrorMiddleware',
 'scrapy.spidermiddlewares.offsite.OffsiteMiddleware',
 'scrapy.spidermiddlewares.referer.RefererMiddleware',
 'scrapy.spidermiddlewares.urllength.UrlLengthMiddleware',
 'scrapy.spidermiddlewares.depth.DepthMiddleware']
2019-05-14 12:32:12 [scrapy.middleware] INFO: Enabled item pipelines:
[]
2019-05-14 12:32:12 [scrapy.core.engine] INFO: Spider opened
2019-05-14 12:32:12 [scrapy.extensions.logstats] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min)
2019-05-14 12:32:12 [scrapy.extensions.telnet] DEBUG: Telnet console listening on 127.0.0.1:6024
2019-05-14 12:32:16 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/robots.txt> (referer: None)
2019-05-14 12:32:17 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html> (referer: None)
2019-05-14 12:32:17 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d15025732-Reviews-Impressive_Resort_Spa_Punta_Cana-Bavaro_Punta_Cana_La_Altagracia_Province_Dominican_.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:18 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d313884-Reviews-Punta_Cana_Princess_All_Suites_Resort_Spa-Bavaro_Punta_Cana_La_Altagracia_Province_Dom.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:18 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d4451011-Reviews-The_Westin_Puntacana_Resort_Club-Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:18 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d10175054-Reviews-Secrets_Cap_Cana_Resort_Spa-Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:18 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d7307251-Reviews-The_Level_at_Melia_Caribe_Beach-Bavaro_Punta_Cana_La_Altagracia_Province_Dominican_Re.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:18 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d292158-Reviews-Grand_Palladium_Punta_Cana_Resort_Spa-Punta_Cana_La_Altagracia_Province_Dominican_Repub.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:18 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d1604057-Reviews-Secrets_Royal_Beach_Punta_Cana-Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:18 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d649099-Reviews-Zoetry_Agua_Punta_Cana-Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:18 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d150842-Reviews-Iberostar_Dominicana_Hotel-Bavaro_Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:19 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d15515013-Reviews-Grand_Memories_Punta_Cana-Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:19 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d150841-Reviews-Iberostar_Selection_Bavaro-Bavaro_Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:19 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d1233228-Reviews-Iberostar_Grand_Bavaro-Bavaro_Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:19 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d149397-Reviews-Bavaro_Princess_Resort_Spa_Casino-Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:19 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d584407-Reviews-Ocean_Blue_Sand-Bavaro_Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:19 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d259337-Reviews-Grand_Palladium_Bavaro_Suites_Resort_Spa-Bavaro_Punta_Cana_La_Altagracia_Province_Domi.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:19 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d150854-Reviews-Hotel_Riu_Palace_Macao-Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:20 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d11701188-Reviews-BlueBay_Grand_Punta_Cana-Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:20 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d14838260-Reviews-Melia_Punta_Cana_Beach_Resort-Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:20 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d1595124-Reviews-Luxury_Bahia_Principe_Esmeralda-Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:20 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d508162-Reviews-Dreams_Punta_Cana_Resort_Spa-Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:20 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d1076311-Reviews-Hard_Rock_Hotel_Casino_Punta_Cana-Bavaro_Punta_Cana_La_Altagracia_Province_Dominican_.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:20 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d10595200-Reviews-Grand_Bahia_Principe_Aquamarine-Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:20 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d611114-Reviews-Hotel_Riu_Palace_Punta_Cana-Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:20 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g3200043-d8709413-Reviews-Excellence_El_Carmen-Uvero_Alto_Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:21 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d1199681-Reviews-Luxury_Bahia_Principe_Ambar-Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:21 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d1889895-Reviews-Karibo_Punta_Cana-Bavaro_Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:21 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d6454132-Reviews-Premium_Level_at_Barcelo_Bavaro_Palace-Bavaro_Punta_Cana_La_Altagracia_Province_Domin.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:21 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d15080584-Reviews-Impressive_Premium_Resorts_Spa-Bavaro_Punta_Cana_La_Altagracia_Province_Dominican_Re.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:21 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g147293-d2687221-Reviews-NH_Punta_Cana-Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:21 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.tripadvisor.com/Hotel_Review-g3176298-d579774-Reviews-Iberostar_Punta_Cana-Bavaro_Punta_Cana_La_Altagracia_Province_Dominican_Republic.html> (referer: https://www.tripadvisor.com/Hotels-g147293-Punta_Cana_La_Altagracia_Province_Dominican_Republic-Hotels.html)
2019-05-14 12:32:21 [scrapy.core.engine] INFO: Closing spider (finished)
2019-05-14 12:32:21 [scrapy.statscollectors] INFO: Dumping Scrapy stats:
{'downloader/request_bytes': 46023,
 'downloader/request_count': 32,
 'downloader/request_method_count/GET': 32,
 'downloader/response_bytes': 5599418,
 'downloader/response_count': 32,
 'downloader/response_status_count/200': 32,
 'finish_reason': 'finished',
 'finish_time': datetime.datetime(2019, 5, 14, 19, 32, 21, 637712),
 'log_count/DEBUG': 33,
 'log_count/INFO': 8,
 'memusage/max': 51412992,
 'memusage/startup': 51412992,
 'request_depth_max': 1,
 'response_received_count': 32,
 'scheduler/dequeued': 31,
 'scheduler/dequeued/memory': 31,
 'scheduler/enqueued': 31,
 'scheduler/enqueued/memory': 31,
 'start_time': datetime.datetime(2019, 5, 14, 19, 32, 12, 996979)}
2019-05-14 12:32:21 [scrapy.core.engine] INFO: Spider closed (finished)

解决方案

This selector is not working:

response.xpath('//div[@class="hotels-review-list-parts-ReviewTitle__reviewTitleText"]/a/@href')

The element on the site is <a> and not <div>, also the class name seems wrong. Perhaps this site append some random data to the class name, as you can see below

<a href="/ShowUserReviews-g3176298-d259337-r673990694-Grand_Palladium_Bavaro_Suites_Resort_Spa-Bavaro_Punta_Cana_La_Altagracia_Provinc.html" class="hotels-review-list-parts-ReviewTitle__reviewTitleText--3QrTy"><span><span>Ótima experiência! Resort amplo, com diversas opções de entretenimento!!!</span></span></a>

You can try to match only part of the string, for example:

//a[contains(@class, "hotels-review-list-parts-ReviewTitle__reviewTitleText")]

这篇关于Tripadvisor 的 Scrapy Spider 抓取了 0 页(以 0 页/分钟的速度)的文章就介绍到这了,希望我们推荐的答案对大家有所帮助,也希望大家多多支持IT屋!

查看全文
登录 关闭
扫码关注1秒登录
发送“验证码”获取 | 15天全站免登陆