why not in python 2.4.3

Collapse
This topic is closed.
X
X
 
  • Time
  • Show
Clear All
new posts
  • Rocco

    #1

    why not in python 2.4.3

    hi
    I made the upgrade to python 2.4.3 from 2.4.2.
    I want to take from google news some atom feeds with a funtion like
    this
    import urllib2
    def takefeed(url):
    request=urllib2 .Request(url)
    request.add_hea der('User-Agent', 'Mozilla/4.0 (compatible; MSIE 5.5;
    Windows NT')
    opener = urllib2.build_o pener()
    data=opener.ope n(request).read ()
    return data
    url='http://news.google.it/?output=rss'
    d=takefeed(url)
    This woks well with python 2.3.5 but does not work with 2.4.3.
    Why?
    Thanks

  • Carl Banks

    #2
    Re: why not in python 2.4.3


    Rocco wrote:[color=blue]
    > hi
    > I made the upgrade to python 2.4.3 from 2.4.2.
    > I want to take from google news some atom feeds with a funtion like
    > this
    > import urllib2
    > def takefeed(url):
    > request=urllib2 .Request(url)
    > request.add_hea der('User-Agent', 'Mozilla/4.0 (compatible; MSIE 5.5;
    > Windows NT')
    > opener = urllib2.build_o pener()
    > data=opener.ope n(request).read ()
    > return data
    > url='http://news.google.it/?output=rss'
    > d=takefeed(url)
    > This woks well with python 2.3.5 but does not work with 2.4.3.
    > Why?[/color]

    Define "woks [sic] well". It works fine for me on 2.4.3 (and by "works
    fine" I mean it ran without an exception and it returned what appeared
    to be RSS data). If you would give us an exception trace it would help
    a lot.

    Maybe Google's server (or your ISP's) was down. That happens
    sometimes.

    Carl

    Comment

    • Rene Pijlman

      #3
      Re: why not in python 2.4.3

      Rocco:[color=blue]
      >but does not work with 2.4.3.[/color]

      Define "does not work".

      --
      René Pijlman

      Comment

      • Rocco

        #4
        Re: why not in python 2.4.3

        This is the problem when I run the function
        this is the result from 2.3.5[color=blue][color=green][color=darkred]
        >>> print rss[/color][/color][/color]
        <?xml version="1.0" encoding="UTF-8"?><feed version="0.3" xml:lang="it"
        xmlns="http://purl.org/atom/ns#"><generator >NFE/1.0</generator><titl e>Google
        News Italia</title><link rel="alternate" type="text/html"
        href="http://news.google.it/"/><tagline>Googl e News
        Italia</tagline><author ><name>Google
        Inc.</name><email>new s-feedback@google .com</email></author><copyrig ht>&amp;copy;20 06
        Google</copyright><modi fied>2006-05-28T19:09:13+00: 00</modified>
        <!-- A couple notes:
        * add an "output=ato m" param to get Atom
        * section pages have a "topic=?" param;
        use "topic=h" for a Top Stories section.
        --><entry><title> Benedetto XVI: Wojtyla santo subito - LibertÃ
        </title><link rel="alternate" type="text/html"
        href="http://www.liberta.it/default.asp?IDG =605282024"/><id>tag:news.g oogle.com,2005: cluster=41b535f b</id><summary>Pri ma
        pagina</summary><issued >2006-05-28T11:05:00+00: 00</issued><modifie d>2006-05-28T11:05:00+00: 00</modified><conte nt
        type="text/html" mode="escaped"> &lt;br&gt;&lt;t able border=0 align=
        cellpadding=5 cellspacing=0&g t;&lt;tr&gt;&lt ;td width=80 align=center
        valign=top&gt;& lt;a .....
        [color=blue][color=green][color=darkred]
        >>> import sys
        >>> sys.getdefaulte ncoding()[/color][/color][/color]
        'ascii'[color=blue][color=green][color=darkred]
        >>>[/color][/color][/color]
        this is the result with 2.4.3[color=blue][color=green][color=darkred]
        >>> print rss[/color][/color][/color]
        ヒ[color=blue][color=green][color=darkred]
        >>> rss[/color][/color][/color]
        '\x1f\x8b\x08\x 00\x00\x00\x00\ x00\x02\xff\xe5 }Ks\xe3F\xb6\xe 6\xfeF\xdc\xff\ x90\xd77\xba\xc 3\x9e\x10D\xbc\ x01\xcaU\xee\xa 1\x9eM[\xa2\xd4$\xabl\ xf7\x86\x93\x04 \x93Tv\x81H\x1a \x0fV\xa9V\xfe\ x0f3\x9b\x8e\x9 8\x89\xb8\xcb\x 1b\xd1\xb3\x9a\ xddD\xef\xec\x7 f\xe2_2\xe7$\x0 0\x8a/\x11|\x93\xd6\x b4\xa3U"\x04\x0 2\x99\xe7d\x9e< \xdfy\xbe\xf9\x d3\xa7\xbeO\x86 ,\x8c\xb8\x08\x de~\xa1\x9d\xaa _\x10\x16x\xa2\ xc3\x83\xde\xdb/\xde5\xaf\x15\x f7\x8b?}\xf3&\x 8c\xa2\xe7\x9bt \xb8\xe9\x9b7\x de#\r\x02\xe6\x 7f\xf3\xa6\xc7\ x02\x16\xd2X\x8 4\xdf\xd4\xae\x afJ\xf0\x847\xa 5\xe7Kob\x1e\xf b\xec\x9b\x1b!z >#5\xf61"\xd5\x 98\xfa\x9c\xbe) \xa5\x7fy\xe3\x f3\xe0\xc37\x8f q<8+\x95\x02\xf 8\xfbiO\xde{\xc a\xe3\xd2\x9b\x 92\xfc\xe3\x9b\ x0e\x8b\xbc\x90 \x0fbx\xfb\xdc\ '\x8d\xff\xfd\x 8dO\x83^B{\xec\ x1b\x1e\xc3\xf7 \xf3\x0fo>\xb2\ xf6\x1d\x8db\x1 6~\x83/Q\xba\x8cu\xda\ xd4\xfb\xf0_\xb 3\xb7y\xa2\xff\ xa6\xf4|\xcf\x1 bO\x0c\x9eB\xde {\x8c\xbf\xf9#\ xed\x0f\xbe\xc6 \x8f_\xeb\xaaj\ x93\xf4\xfdoJ\x cf7\xbc\x19$\xe dK\x1a\xb3o\x1a IpBt\x97\xdc\xd 1\'"\xef\xd5\xb 53\xcd<3\x1crs\ xd7|S\xcao\x83\ x11F\xf1y\xc2\x fd\xce2\xdf\x9a \xbc\xf9_\xff\x e5\xcd\xbf)\n\x a9\x10O$\x03
        C b\x16\x9d\xfd\x eb\xbf\x10\xfc\ xdf\x7f!\xb4\xd 3!4
        _\x88$\x1e$\xf1[\xe0\xda\x17d@C \xda\'\xb1
        =\x16\x93z\xa31 \xba9b\x1eR\x0c n\xe8\xb1\x88<\ xd2!#\x94|\x11\ x8b\x01\xf7\xde \xfe)\xfb\xde\x d7\xd9\xdd\x84$ \x11\xcb\xff\xf 8\xf8\x05\xe9\x 8a\x10nn\x8a\x0 1i\x00\x979|?{\ xda\xe9\xbf\xfe \x8b\xa2|\xf3\x 86\xf7%\xd1\x0b \x99\x9f\x84\xf e<\xde\x037J<\x 88\xfd\x12\x8f[\xb0\x0e\xe4\xd 3"yG+dp\x17\xef \xbe)\xe1W\x97Y <\xa5l,<f\xfd|D \x15\xf2=\xed\x 88\x8f\xdcc\'\x c4\xe3q\xfc\xeb \x7f\x90\x80\xc 2\xc8\x18\xe9p\ xf2\x1d\r\x85\x 7fB\xb8O\x1eD\x 10\xb3.\xdc\x05 \xd4!\xb4\xdbea \x1f\x1659==%\n \xa9\xfa\xa4\xc 9\xfa\x03\xb1\x ccB\xc6\x8f8\xe 0?E\xf4m3]P\xf1[\xb8\xae*\xf0\x 9f\xfc\xdc\xed\ xbc\xad\xcb_\xe 0\xae\xb7\xd9C> ~\xfcx\xca\xfd\ x18_\x82\x0f\xa 1\x83A(\xba"\xe 8\xf0>\x0bb\x0e \x04\xea\xb0O\x a74\x
        [color=blue][color=green][color=darkred]
        >>> import sys
        >>> sys.getdefaulte ncoding()[/color][/color][/color]
        'latin_1'[color=blue][color=green][color=darkred]
        >>>[/color][/color][/color]
        No exception trace
        Thanks again

        Comment

        • Serge Orlov

          #5
          Re: why not in python 2.4.3

          Rocco wrote:
          [color=blue][color=green][color=darkred]
          > >>> import sys
          > >>> sys.getdefaulte ncoding()[/color][/color]
          > 'latin_1'[/color]

          Don't change default encoding. It should be always ascii.

          Comment

          • Rocco

            #6
            Re: why not in python 2.4.3

            Also with ascii the function does not work.

            Comment

            • Serge Orlov

              #7
              Re: why not in python 2.4.3

              Rocco wrote:[color=blue]
              > Also with ascii the function does not work.[/color]

              Well, at least you fixed misconfiguratio n ;)

              Googling for 1F8B (that's two first bytes from your strange python 2.4
              result) gives a hint: it's a beginning of gzip stream. Maybe urllib2 in
              python 2.4 reports to the server that it supports compressed data but
              doesn't decompress it when receives the reply?

              Comment

              • Rocco

                #8
                Re: why not in python 2.4.3

                Thanks Serge.
                It's a gzip string.
                So the code is[color=blue][color=green][color=darkred]
                >>> import urllib2
                >>> def takefeed(url):[/color][/color][/color]
                request=urllib2 .Request(url)
                request.add_hea der('User-Agent', 'Mozilla/4.0 (compatible; MSIE
                5.5;Windows NT')
                opener = urllib2.build_o pener()
                data=opener.ope n(request).read ()
                return data
                [color=blue][color=green][color=darkred]
                >>> url='http://news.google.it/?output=rss'
                >>> d=takefeed(url)
                >>> from StringIO import StringIO
                >>> zipdata=StringI O(d)
                >>> import gzip
                >>> gz=gzip.GzipFil e(fileobj=zipda ta)
                >>> rss=gz.read()
                >>> len(rss)[/color][/color][/color]
                102529[color=blue][color=green][color=darkred]
                >>> print rss[0:100][/color][/color][/color]
                <?xml version="1.0" encoding="UTF-8"?><rss
                version="2.0">< channel><genera tor>NFE/1.0</generator><tit[color=blue][color=green][color=darkred]
                >>>[/color][/color][/color]

                Comment

                • John Machin

                  #9
                  Re: why not in python 2.4.3

                  On 29/05/2006 10:47 PM, Serge Orlov wrote:[color=blue]
                  > Rocco wrote:[color=green]
                  >> Also with ascii the function does not work.[/color]
                  >
                  > Well, at least you fixed misconfiguratio n ;)
                  >
                  > Googling for 1F8B (that's two first bytes from your strange python 2.4
                  > result) gives a hint: it's a beginning of gzip stream.[/color]

                  Well done!
                  [color=blue]
                  > Maybe urllib2 in
                  > python 2.4 reports to the server that it supports compressed data but
                  > doesn't decompress it when receives the reply?
                  >[/color]

                  Something funny is happening here. Others reported it working with 2.4.3
                  and Rocco's original code as posted in this thread -- which works for me
                  on 2.4.2, Windows XP.

                  There was one suss thing about Rocco's problem description:
                  First message ended with d=takefeed(url)
                  But next message said print rss
                  Is rss == d?

                  Cheers,
                  John

                  Comment

                  • John Machin

                    #10
                    Re: why not in python 2.4.3

                    On 30/05/2006 12:44 AM, Rocco wrote:[color=blue]
                    > Thanks Serge.
                    > It's a gzip string.[/color]

                    Look, Ma, no gzip!!!

                    C:\junk>rocco_r ss.py
                    '<?xml version="1.0" encoding="UTF-8"?><rss
                    version="2.0">< channel><genera tor>NF
                    E/1.0</generator><tit'

                    C:\junk>type rocco_rss.py
                    import urllib2
                    def takefeed(url):
                    request=urllib2 .Request(url)
                    request.add_hea der('User-Agent', 'Mozilla/4.0 (compatible; MSIE
                    5.5; Win
                    dows NT')
                    opener = urllib2.build_o pener()
                    data=opener.ope n(request).read ()
                    return data
                    url='http://news.google.it/?output=rss'
                    d=takefeed(url)
                    print repr(d[:100])

                    Comment

                    • Serge Orlov

                      #11
                      Re: why not in python 2.4.3

                      John Machin wrote:[color=blue]
                      > On 29/05/2006 10:47 PM, Serge Orlov wrote:[color=green]
                      > > Maybe urllib2 in
                      > > python 2.4 reports to the server that it supports compressed data but
                      > > doesn't decompress it when receives the reply?
                      > >[/color]
                      >
                      > Something funny is happening here. Others reported it working with 2.4.3
                      > and Rocco's original code as posted in this thread -- which works for me
                      > on 2.4.2, Windows XP.[/color]

                      It "works" for me too, returning raw uncompressed data.
                      [color=blue]
                      > There was one suss thing about Rocco's problem description:
                      > First message ended with d=takefeed(url)
                      > But next message said print rss
                      > Is rss == d?[/color]

                      Nope. If you look at html tags, 2.3 code returns <feed> <generator> ...
                      whereas 2.4 code returns <rss> <channel> <generator> ... That may
                      explain why 2.3 result is not compressed and 2.4 result is compressed,
                      but that doesn't explain why 2.4 *is* compressed. I looked at python
                      2.4 httplib, I'm sure it's not a problem, quote from httplib:

                      # we only want a Content-Encoding of "identity" since we
                      don't
                      # support encodings such as x-gzip or x-deflate.

                      I think there is a web accellerator sitting somewhere between Rocco and
                      Google server that is confused that Rocco is "misinformi ng" web server
                      saying he's using Firefox, but at the same time claiming that he cannot
                      handle compressed data. That's why they teach little kids: don't lie :)

                      Comment

                      Working...