elementtidy, \0 chars and parsing from a string

Collapse
This topic is closed.
X
X
 
  • Time
  • Show
Clear All
new posts
  • Steven Bethard

    #1

    elementtidy, \0 chars and parsing from a string

    So I see that elementtidy doesn't like strings with \0 characters in them:
    [color=blue][color=green][color=darkred]
    >>> import urllib
    >>> from elementtidy import TidyHTMLTreeBui lder
    >>> url = 'http://news.bbc.co.uk/1/hi/world/europe/492215.stm'
    >>> url_file = urllib.urlopen( url)
    >>> tree = TidyHTMLTreeBui lder.parse(url_ file)[/color][/color][/color]
    Traceback (most recent call last):
    ...
    File "...elementtidy \TidyHTMLTreeBu ilder.py", line 90, in close
    stdout, stderr = _elementtidy.fi xup(*args)
    TypeError: fixup() argument 1 must be string without null bytes, not str

    The obvious solution would be to str.replace('\0 ', '') on the file's
    text, but I'm not sure how to ask elementtidy to parse from a string
    instead of a file-like object. Do I need to wrap it in a StringIO, or
    is there a better way?

    STeVe
  • Simon Percivall

    #2
    Re: elementtidy, \0 chars and parsing from a string

    Well, it seems you can do:

    parser = elementtidy.Tid yHTMLTreeBuilde r.TidyHTMLTreeB uilder()
    parser.feed(you r_str)
    tree = elementtree.Ele mentTree.Elemen tTree(element=p arser.close())

    Look at the parse() method in the ElementTree class.

    Comment

    Working...