error when parsing xml

Collapse
This topic is closed.
X
X
 
  • Time
  • Show
Clear All
new posts
  • Albert Leibbrandt

    #1

    error when parsing xml

    >> I use xml.dom.minidom to parse some xml, but when input[color=blue][color=green]
    >> contains some specific caracters(æ, ø and å), I get an
    >> UnicodeEncodeEr ror, like this:
    >>
    >> UnicodeEncodeEr ror: 'ascii' codec can't encode character
    >> u'\xe6' in position 604: ordinal not in range(128).
    >>
    >> How can I avoid this error?
    >>
    >>
    >> All help much appreciated![/color][/color]

    I have found that some people refuse to stick to standards, so whenever I
    parse XML files I remove any characters that fall in the range
    <= 0x1f[color=blue]
    >= 0xf0[/color]

    Hope it helps.

    Regards
    Albert


  • Diez B. Roggisch

    #2
    Re: error when parsing xml

    > I have found that some people refuse to stick to standards, so whenever I[color=blue]
    > parse XML files I remove any characters that fall in the range
    > <= 0x1f
    >[color=green]
    >>= 0xf0[/color][/color]

    Now of what help shall that be? Get rid of all accented characters?
    Sorry, but that surely is the dumbest thing to do here - and has
    _nothing_ to do with standards! Charactersets with codepoints > 128 are
    pretty common and well standarized, just not "ascii". I suggset you read
    up on the topic of unicode & encodings a bit - and then fix some code of
    yours...

    Diez

    Comment

    • Richard Brodie

      #3
      Re: error when parsing xml

      [color=blue][color=green]
      > > I have found that some people refuse to stick to standards, so whenever I
      > > parse XML files I remove any characters that fall in the range
      > > <= 0x1f
      > >[color=darkred]
      > >>= 0xf0[/color][/color]
      >
      > Now of what help shall that be? Get rid of all accented characters?
      > Sorry, but that surely is the dumbest thing to do here - and has
      > _nothing_ to do with standards![/color]

      Earlier versions of the Microsoft XML parser accept invalid characters
      (e.g. most of those < 0x1f). Sadly, you do find files in the wild that need
      to have these stripped before feeding them to a conforming parser.
      One can be too enthusiastic about the process, though.


      Comment

      Working...