ascii character - removing chars from string

Collapse
This topic is closed.
X
X
 
  • Time
  • Show
Clear All
new posts
  • bruce

    #1

    ascii character - removing chars from string

    hi...

    i'm running into a problem where i'm seeing non-ascii chars in the parsing
    i'm doing. in looking through various docs, i can't find functions to
    remove/restrict strings to valid ascii chars.

    i'm assuming python has something like

    valid_str = strip(invalid_s tr)

    where 'strip' removes/strips out the invalid chars...

    any ideas/thoughts/pointers...

    thanks

    -bruce

  • bearophileHUGS@lycos.com

    #2
    Re: ascii character - removing chars from string

    bruce:
    valid_str = strip(invalid_s tr)
    where 'strip' removes/strips out the invalid chars...
    This isn't short but it is fast:
    import string
    valid_chars = string.lowercas e + string.uppercas e + \
    string.digits +
    """|!'\\"£$ %&/()=?^*é§_:;>+,.-<\n \t"""
    all_chars = "".join(map ( chr, range(256)) )
    comp_valid_char s = "".join( set(all_chars). difference(vali d_chars) )
    print "test string".transla te(all_chars, comp_valid_char s)


    Shorter and a bit slower alternative:
    import string
    valid_chars_set = set(string.lowe rcase + string.uppercas e
    + string.digits +
    """|!'\\"£$ %&/()=?^*é§_:;>+,.-<\n \t""")
    print filter(lambda c: c in valid_chars_set , "test string")

    You can add the chars you want to the string of accepted ones.

    Bye,
    bearophile

    Comment

    • John Machin

      #3
      Re: ascii character - removing chars from string

      On 4/07/2006 9:27 AM, bruce wrote:
      hi...
      >
      i'm running into a problem where i'm seeing non-ascii chars in the parsing
      i'm doing. in looking through various docs, i can't find functions to
      remove/restrict strings to valid ascii chars.
      >
      It's possible that you would be better off handling those characters in
      some fashion other than blowing them away. What are the characters that
      you are seeing, and what is the problem that they are causing you?

      Comment

      • Rune Strand

        #4
        Re: ascii character - removing chars from string

        bruce wrote:
        hi...
        >
        i'm running into a problem where i'm seeing non-ascii chars in the parsing
        i'm doing. in looking through various docs, i can't find functions to
        remove/restrict strings to valid ascii chars.
        >
        i'm assuming python has something like
        >
        valid_str = strip(invalid_s tr)
        >
        where 'strip' removes/strips out the invalid chars...
        >
        any ideas/thoughts/pointers...
        If you're able to define the invalid_chars, the most convenient is
        probably to use the strip() method:
        >>a_string = "abcdef"
        >>invalid_cha rs = 'abc'
        >>a_string.stri p(invalid_chars )
        'def'

        Comment

        • Simon Forman

          #5
          Re: ascii character - removing chars from string

          bruce wrote:
          hi...
          >
          i'm running into a problem where i'm seeing non-ascii chars in the parsing
          i'm doing. in looking through various docs, i can't find functions to
          remove/restrict strings to valid ascii chars.
          >
          i'm assuming python has something like
          >
          valid_str = strip(invalid_s tr)
          >
          where 'strip' removes/strips out the invalid chars...
          >
          any ideas/thoughts/pointers...
          >
          thanks
          >
          -bruce
          You might be able to use the translate() and maketrans() string
          methods. See

          second-to-last post for an example.

          Comment

          • Simon Forman

            #6
            Re: ascii character - removing chars from string

            bruce wrote:
            hi...
            >
            update. i'm getting back html, and i'm getting strings like " foo &nbsp;"
            which is valid HTML as the '&nbsp;' is a space.
            &, n, b, s, p, ; Those are all ascii characters.
            i need a way of stripping/removing the '&nbsp;' from the string
            >
            the &nbsp; needs to be treated as a single char...
            >
            text = "foo cat &nbsp;"
            >
            ie ok_text = strip(text)
            >
            ok_text = "foo cat"
            Do you really want to remove those html entities? Or would you rather
            convert them back into the actual text they represent? Do you just
            want to deal with &nbsp;'s? Or maybe the other possible entities that
            might appear also?

            Check out htmlentitydefs. entitydefs (see
            http://docs.python.org/lib/module-htmlentitydefs.html) it's kind of
            ugly looking so maybe use pprint to print it:
            >>import htmlentitydefs, pprint
            >>pprint.pprint (htmlentitydefs .entitydefs)
            {'AElig': 'Æ',
            'Aacute': 'Á',
            'Acirc': 'Â',
            ..
            ..
            ..
            'nbsp': '\xa0',
            ..
            ..
            ..
            etc...


            HTH,
            ~Simon

            "You keep using that word. I do not think it means what you think it
            means."
            -Inigo Montoya, "The Princess Bride"

            Comment

            • Simon Forman

              #7
              Re: ascii character - removing chars from string

              bruce wrote:
              simon...
              >
              the '&nbsp;' is not to be seen/viewed as text/ascii.. it's a representation
              of a hex 'u\xa0' if i recall...
              Did you not see this part of the post that you're replying to?
              'nbsp': '\xa0',
              My point was not that '\xa0' is an ascii character... It was that your
              initial request was very misleading:

              "i'm running into a problem where i'm seeing non-ascii chars in the
              parsing i'm doing. in looking through various docs, i can't find
              functions to remove/restrict strings to valid ascii chars."

              That's why you got three different answers to the wrong question.

              You weren't "seeing non-ascii chars" at all. You were seeing ascii
              representations of html entities that, in the case of '&nbsp;', happen
              to represent non-ascii values.
              >
              i'm looking to remove or replace the insances with a ' ' (space)
              Simplicity:

              s.replace('&nbs p;', ' ')

              ~Simon

              "You keep using that word. I do not think it means what you think it
              means."
              -Inigo Montoya, "The Princess Bride"
              >
              -bruce
              >
              >
              -----Original Message-----
              From: python-list-bounces+bedougl as=earthlink.ne t@python.org
              [mailto:python-list-bounces+bedougl as=earthlink.ne t@python.org]On Behalf
              Of Simon Forman
              Sent: Monday, July 03, 2006 7:17 PM
              To: python-list@python.org
              Subject: Re: ascii character - removing chars from string
              >
              >
              bruce wrote:
              hi...

              update. i'm getting back html, and i'm getting strings like " foo &nbsp;"
              which is valid HTML as the '&nbsp;' is a space.
              >
              &, n, b, s, p, ; Those are all ascii characters.
              >
              i need a way of stripping/removing the '&nbsp;' from the string

              the &nbsp; needs to be treated as a single char...

              text = "foo cat &nbsp;"

              ie ok_text = strip(text)

              ok_text = "foo cat"
              >
              Do you really want to remove those html entities? Or would you rather
              convert them back into the actual text they represent? Do you just
              want to deal with &nbsp;'s? Or maybe the other possible entities that
              might appear also?
              >
              Check out htmlentitydefs. entitydefs (see
              http://docs.python.org/lib/module-htmlentitydefs.html) it's kind of
              ugly looking so maybe use pprint to print it:
              >
              >import htmlentitydefs, pprint
              >pprint.pprint( htmlentitydefs. entitydefs)
              {'AElig': 'Æ',
              'Aacute': 'Á',
              'Acirc': 'Â',
              .
              .
              .
              'nbsp': '\xa0',
              .
              .
              .
              etc...
              >
              >
              HTH,
              ~Simon
              >
              "You keep using that word. I do not think it means what you think it
              means."
              -Inigo Montoya, "The Princess Bride"

              --
              http://mail.python.org/mailman/listinfo/python-list

              Comment

              Working...