解析非标准XML(CDATA标记) [英] Parsing non-standard XML (CDATA tag)

查看：177 发布时间：2020/9/20 6:08:26 python xml beautifulsoup

本文介绍了解析非标准XML(CDATA标记)的处理方法，对大家解决问题具有一定的参考价值，需要的朋友们下面随着小编来一起学习吧！

问题描述

当我想使用BeautifulSoup库在Python中解析XML文档时，我遇到了一些问题.我要解析的XML文档:

When I want to parsing XML document in Python using BeautifulSoup library, I faced some problems. The XML document that I want to parse:

<item>
<title><![CDATA[Title Sample]]></title>
<link /><![CDATA[http://banhada.kr/?cateCode=09&viewCode=S0941580]]>
<time_start>2011-10-10 09:00:00</time_start>
<time_end>2011-10-17 09:00:00</time_end>
<price_original>35000</price_original>
<price_now>20000</price_now>
</item>

正如您在上面看到的，tag有点奇怪.在我看来，that(标签)不是标准的XML格式，对吗?我如何解析这种可怕的形式?

As you can see above, tag is a little strange. In my opinion, that( tag) is not a stand XML form, right? How can I parse this terrible form?

推荐答案

您不需要BeautifulStoneSoup或lxml. Python附带的电池可以很好地完成这项工作，并且似乎没有任何与XML不兼容的东西.

You don't need BeautifulStoneSoup or lxml. Python's included batteries do the job just fine, and there doesn't seem to be anything non-compliant about your XML.

>>> content='''\
... <item>
... <title><![CDATA[Title Sample]]></title>
... <link /><![CDATA[http://banhada.kr/?cateCode=09&viewCode=S0941580]]>
... <time_start>2011-10-10 09:00:00</time_start>
... <time_end>2011-10-17 09:00:00</time_end>
... <price_original>35000</price_original>
... <price_now>20000</price_now>
... </item>'''
>>> import xml.etree.cElementTree as et
>>> foo = et.XML(content)
>>> for e in foo:
...     print e.tag, e.text, repr(e.tail)
...
title Title Sample '\n'
link None 'http://banhada.kr/?cateCode=09&viewCode=S0941580\n'
time_start 2011-10-10 09:00:00 '\n'
time_end 2011-10-17 09:00:00 '\n'
price_original 35000 '\n'
price_now 20000 '\n'
>>>

这篇关于解析非标准XML(CDATA标记)的文章就介绍到这了，希望我们推荐的答案对大家有所帮助，也希望大家多多支持IT屋！

查看全文

解析非标准XML(CDATA标记) [英] Parsing non-standard XML (CDATA tag)

问题描述

推荐答案

相关文章

Python最新文章

热门教程

热门工具

登录关闭

解析非标准XML(CDATA标记) [英] Parsing non-standard XML (CDATA tag)

问题描述

推荐答案

相关文章

Python最新文章

热门教程

热门工具

登录 关闭

登录关闭