Scrapy是否有可能从原始HTML数据中获取纯文本? [英] Is it possible for Scrapy to get plain text from raw HTML data?

查看：87 发布时间：2020/11/24 3:50:06 python html web-scraping scrapy web-crawler

本文介绍了Scrapy是否有可能从原始HTML数据中获取纯文本?的处理方法，对大家解决问题具有一定的参考价值，需要的朋友们下面随着小编来一起学习吧！

问题描述

例如:

scrapy shell http://scrapy.org/
content = hxs.select('//*[@id="content"]').extract()[0]
print content

然后，我得到以下原始HTML代码:

Then, I get the following raw HTML code:

<div id="content">


  <h2>Welcome to Scrapy</h2>

  <h3>What is Scrapy?</h3>

  <p>Scrapy is a fast high-level screen scraping and web crawling
    framework, used to crawl websites and extract structured data from their
    pages. It can be used for a wide range of purposes, from data mining to
    monitoring and automated testing.</p>

  <h3>Features</h3>

  <dl>

    <dt>Simple</dt>
    <dt>
    </dt>
    <dd>Scrapy was designed with simplicity in mind, by providing the features
      you need without getting in your way
    </dd>

    <dt>Productive</dt>
    <dd>Just write the rules to extract the data from web pages and let Scrapy
      crawl the entire web site for you
    </dd>

    <dt>Fast</dt>
    <dd>Scrapy is used in production crawlers to completely scrape more than
      500 retailer sites daily, all in one server
    </dd>

    <dt>Extensible</dt>
    <dd>Scrapy was designed with extensibility in mind and so it provides
      several mechanisms to plug new code without having to touch the framework
      core

    </dd>
    <dt>Portable, open-source, 100% Python</dt>
    <dd>Scrapy is completely written in Python and runs on Linux, Windows, Mac and BSD</dd>

    <dt>Batteries included</dt>
    <dd>Scrapy comes with lots of functionality built in. Check <a
        href="http://doc.scrapy.org/en/latest/intro/overview.html#what-else">this
      section</a> of the documentation for a list of them.
    </dd>

    <dt>Well-documented &amp; well-tested</dt>
    <dd>Scrapy is <a href="/doc/">extensively documented</a> and has an comprehensive test suite
      with <a href="http://static.scrapy.org/coverage-report/">very good code
        coverage</a></dd>

    <dt><a href="/community">Healthy community</a></dt>
    <dd>
      1,500 watchers, 350 forks on Github (<a href="https://github.com/scrapy/scrapy">link</a>)<br>
      700 followers on Twitter (<a href="http://twitter.com/ScrapyProject">link</a>)<br>
      850 questions on StackOverflow (<a href="http://stackoverflow.com/tags/scrapy/info">link</a>)<br>
      200 messages per month on mailing list (<a
        href="https://groups.google.com/forum/?fromgroups#!aboutgroup/scrapy-users">link</a>)<br>
      40-50 users always connected to IRC channel (<a href="http://webchat.freenode.net/?channels=scrapy">link</a>)
    </dd>

    <dt><a href="/support">Commercial support</a></dt>
    <dd>A few companies provide Scrapy consulting and support</dd>

    <p>Still not sure if Scrapy is what you're looking for?. Check out <a
        href="http://doc.scrapy.org/en/latest/intro/overview.html">Scrapy at a
      glance</a>.

    </p>
    <h3>Companies using Scrapy</h3>

    <p>Scrapy is being used in large production environments, to crawl
      thousands of sites daily. Here is a list of <a href="/companies/">Companies
        using Scrapy</a>.</p>

    <h3>Where to start?</h3>

    <p>Start by reading <a href="http://doc.scrapy.org/en/latest/intro/overview.html">Scrapy at a glance</a>,
      then <a href="/download/">download Scrapy</a> and follow the <a
          href="http://doc.scrapy.org/en/latest/intro/tutorial.html">Tutorial</a>.


    </p></dl>
</div>

但是我想直接从scrapy中获取纯文本.

But I want to get plain text directly from scrapy.

我不想使用任何xPath选择器来提取p，h2，h3 ...标签，因为我正在抓取一个主要内容嵌入到table，;递归地找到xPath可能是一项繁琐的任务.

I do not want to use any xPath selectors to extract the p, h2, h3... tags, since I am crawling a website whose main content is embedded into a table, tbody; recursively. It can be a tedious task to find the xPath.

这可以通过Scrapy中的内置功能实现吗?还是我需要外部工具对其进行转换?我已经阅读了Scrapy的所有文档，但一无所获.

Can this be implemented by a built in function in Scrapy? Or do I need external tools to convert it? I have read through all of Scrapy's docs, but have gained nothing.

这是一个示例站点，可以将原始HTML转换为纯文本: http://beaker .mailchimp.com/html-to-text

This is a sample site which can convert raw HTML into plain text: http://beaker.mailchimp.com/html-to-text

推荐答案

Scrapy没有内置的此类功能.您正在寻找 html2text .

Scrapy doesn't have such functionality built-in. html2text is what you are looking for.

这里是一个示例蜘蛛，它抓取了维基百科的python页面，并使用xpath并使用html2text将html转换为纯文本:

Here's a sample spider that scrapes wikipedia's python page, gets first paragraph using xpath and converts html into plain text using html2text:

from scrapy.selector import HtmlXPathSelector
from scrapy.spider import BaseSpider
import html2text


class WikiSpider(BaseSpider):
    name = "wiki_spider"
    allowed_domains = ["www.wikipedia.org"]
    start_urls = ["http://en.wikipedia.org/wiki/Python_(programming_language)"]

    def parse(self, response):
        hxs = HtmlXPathSelector(response)
        sample = hxs.select("//div[@id='mw-content-text']/p[1]").extract()[0]

        converter = html2text.HTML2Text()
        converter.ignore_links = True
        print(converter.handle(sample)) #Python 3 print syntax

打印:

** Python **是一种广泛使用的通用高级编程语言.[11] [12] [13]它的设计理念强调代码可读性及其语法使程序员可以用代码行比诸如 C. [14] [15]该语言提供了旨在使清除变得清晰的结构. 程序的规模不限.[16]

**Python** is a widely used general-purpose, high-level programming language.[11][12][13] Its design philosophy emphasizes code readability, and its syntax allows programmers to express concepts in fewer lines of code than would be possible in languages such as C.[14][15] The language provides constructs intended to enable clear programs on both a small and large scale.[16]

这篇关于Scrapy是否有可能从原始HTML数据中获取纯文本?的文章就介绍到这了，希望我们推荐的答案对大家有所帮助，也希望大家多多支持IT屋！

查看全文

Scrapy是否有可能从原始HTML数据中获取纯文本? [英] Is it possible for Scrapy to get plain text from raw HTML data?

问题描述

推荐答案

相关文章

前端开发最新文章

热门教程

热门工具

登录关闭

Scrapy是否有可能从原始HTML数据中获取纯文本? [英] Is it possible for Scrapy to get plain text from raw HTML data?

问题描述

推荐答案

相关文章

前端开发最新文章

热门教程

热门工具

登录 关闭

登录关闭