jsoup 1.6.2發(fā)布最棒的Java HTML解析器

作者：傾楓 2012-03-29 11:48:28

jsoup 是一款 Java 的HTML 解析器，可直接解析某個(gè)URL地址、HTML文本內(nèi)容。它提供了一套非常省力的API，可通過DOM，CSS以及類似于JQuery的操作方法來取出和操作數(shù)據(jù)。

jsoup 1.6.2 發(fā)布了，改版包含很多的 bug 修復(fù)，松散的 XML 解析模式，功能調(diào)整以及內(nèi)存的改進(jìn)。

主要改進(jìn)內(nèi)容包括：

- Added a simplified XML parsing mode, which can usefully parse valid and invalid XML, but does not enforce any HTML document structure or special tag behaviour.
- Added the optional ability to track errors when tokenising and parsing.
- Added Jsoup.connect.cookies(Map) method, to set multiple cookies at once, possibly from a prior request.
- Added Element.textNodes() and Element.dataNodes(), to easily access an element's children text nodes and data nodes.
- Added an example program that demonstrates how to format HTML as plain-text, and the use of the NodeVisitor interface.
- Added Node.traverse() and Elements.traverse() methods, to iterate through a node's descendants.
- Updated Jsoup.connect() so that when requests made as POSTs are redirected, the redirect is followed as a GET.
- Updated the Cleaner and whitelists to optionally preserve related links in elements, instead of converting them to absolute links.
- Updated the Cleaner to support custom allowed protocols such as "cid:" and "data:".
- Updated handling of base href tags, to act on only the first one seen when parsing, to align with modern browsers.
- Updated Node.setBaseUri(), to recursively set on all the node's descendants.
Bug fixes:
- Fixed an issue where all HTML parse errors where being tracked as new objects, creating high memory pressure on low-memory devices.
- Fixed handling of null characters within comments.
- Tweaked escaped entity detection in attributes to not treat &entity_... as an entity form.
- Fixed doctype tokeniser to allow whitespace between name and public identifier.
- Fixed issue where comments within a table tag would be duplicate-fostered into body.
- Fixed an issue where a spurious byte-order-mark at the start of a document would cause the parser to miss head contents.
- Fixed an issue where content after a frameset could cause a NPE crash. Now correctly implements spec and ignores the trailing content.
- Tweaked whitespace checks to align with HTML spec.
- Tweaked HTML output of closing script and style tags to not add an extraneous newline when pretty-printing.
- Substantially reduced default memory allocation within Node.outerHtml, to reduce memory pressure when serialising smaller DOMs.

詳情請看官方發(fā)行說明：

http://jsoup.org/news/release-1.6.2

jsoup的主要功能如下：

從一個(gè)URL，文件或字符串中解析HTML；
使用DOM或CSS選擇器來查找、取出數(shù)據(jù)；
可操作HTML元素、屬性、文本；

jsoup是基于MIT協(xié)議發(fā)布的，可放心使用于商業(yè)項(xiàng)目。

示例代碼：

File input = new File("/tmp/input.html");  
Document doc = Jsoup.parse(input, "UTF-8", "http://example.com/");  
 
Element content = doc.getElementById("content");  
Elements links = content.getElementsByTag("a");  
for (Element link : links) {  
  String linkHref = link.attr("href");  
  String linkText = link.text();  
}

下載地址：http://jsoup.org/download

【編輯推薦】

JActor 2.2.0 RC3發(fā)布 Actor模式的Java實(shí)現(xiàn)
LogicalDOC 6.4發(fā)布 Java開源文檔管理系統(tǒng)
Resin 4.0.27發(fā)布 Java應(yīng)用服務(wù)器
LibrePlan 1.2.2發(fā)布 Java開源項(xiàng)目計(jì)劃和管理
xmemcached 1.3.6發(fā)布 memcached的Java開發(fā)包

責(zé)任編輯：林師授來源： 51CTO

jsoup Java

自拍偷在线精品自拍偷,亚洲欧美中文日韩v在线观看不卡

jsoup 1.6.2發(fā)布 最棒的Java HTML解析器

jsoup 1.6.2發(fā)布最棒的Java HTML解析器