Making HTTP requests using Twisted

Collapse
This topic is closed.
X
X
 
  • Time
  • Show
Clear All
new posts
  • rzimerman

    #1

    Making HTTP requests using Twisted

    I'm hoping to write a program that will read any number of urls from
    stdin (1 per line), download them, and process them. So far my script
    (below) works well for small numbers of urls. However, it does not
    scale to more than 200 urls or so, because it issues HTTP requests for
    all of the urls simultaneously, and terminates after 25 seconds.
    Ideally, I'd like this script to download at most 50 pages in parallel,
    and to time out if and only if any HTTP request is not answered in 3
    seconds. What changes do I need to make?

    Is Twisted the best library for me to be using? I do like Twisted, but
    it seems more suited to batch mode operations. Is there some way that I
    could continue registering url requests while the reactor is running?
    Is there a way to specify a time out per page request, rather than for
    a batch of pages requests?

    Thanks!



    #-------------------------------------------------

    from twisted.interne t import reactor
    from twisted.web import client
    import re, urllib, sys, time

    def extract(html):
    #do some processing on html, writing to stdout

    def printError(fail ure):
    print >sys.stderr, "Error:", failure.getErro rMessage( )

    def stopReactor():
    print "Now stopping reactor..."
    reactor.stop()

    for url in sys.stdin:
    url = url.rstrip()
    client.getPage( url).addCallbac k(extract).addE rrback(printErr or)

    reactor.callLat er(25, stopReactor)
    reactor.run()

  • K.S.Sreeram

    #2
    Re: Making HTTP requests using Twisted

    rzimerman wrote:
    I'm hoping to write a program that will read any number of urls from
    stdin (1 per line), download them, and process them. So far my script
    (below) works well for small numbers of urls. However, it does not
    scale to more than 200 urls or so, because it issues HTTP requests for
    all of the urls simultaneously, and terminates after 25 seconds.
    Ideally, I'd like this script to download at most 50 pages in parallel,
    and to time out if and only if any HTTP request is not answered in 3
    seconds. What changes do I need to make?

    Is Twisted the best library for me to be using? I do like Twisted, but
    it seems more suited to batch mode operations. Is there some way that I
    could continue registering url requests while the reactor is running?
    Is there a way to specify a time out per page request, rather than for
    a batch of pages requests?
    Have a look at pyCurl. (http://pycurl.sourceforge.net)

    Regards
    Sreeram



    -----BEGIN PGP SIGNATURE-----
    Version: GnuPG v1.4.2.2 (MingW32)
    Comment: Using GnuPG with Mozilla - http://enigmail.mozdev.org

    iD8DBQFEs2vIrgn 0plK5qqURAmahAJ 4oPAJ4AtPNvRFxs 99IFNHuViyCiQCg mT8a
    GYqpz82zvsin4Qr XGXW0WDI=
    =rz4Q
    -----END PGP SIGNATURE-----

    Comment

    • Fredrik Lundh

      #3
      Re: Making HTTP requests using Twisted

      "rzimerman" wrote:
      Is Twisted the best library for me to be using? I do like Twisted, but
      it seems more suited to batch mode operations. Is there some way that I
      could continue registering url requests while the reactor is running?
      Is there a way to specify a time out per page request, rather than for
      a batch of pages requests?
      there are probably ways to solve this with Twisted, but in case you want a
      simpler alternative, you could use Python's standard asyncore module and
      the stuff described here:



      especially




      </F>



      Comment

      • Manlio Perillo

        #4
        Re: Making HTTP requests using Twisted

        rzimerman ha scritto:
        I'm hoping to write a program that will read any number of urls from
        stdin (1 per line), download them, and process them. So far my script
        (below) works well for small numbers of urls. However, it does not
        scale to more than 200 urls or so, because it issues HTTP requests for
        all of the urls simultaneously, and terminates after 25 seconds.
        Ideally, I'd like this script to download at most 50 pages in parallel,
        and to time out if and only if any HTTP request is not answered in 3
        seconds. What changes do I need to make?
        >
        Take a look at


        And read


        You can pass a timeout to the constructor.

        To download at most 50 pages in parallel you can use a download queue.

        Here is a quick example, ABSOLUTELY NOT TESTED:

        class DownloadQueue(o bject):
        SIZE = 50

        def init(self):
        self.requests = [] # queued requests
        self.deferreds = [] # waiting requests

        def addRequest(self , url, timeout):
        if len(self.deferr eds) >= sels.SIZE:
        # wait for completion of all previous requests
        DeferredList(se lf.deferreds
        ).addCallback(s elf._callback)
        self.deferreds = []

        # queue the request
        deferred = Deferred()
        self.requests.a ppend((url, timeout, deferred))

        return deferred
        else:
        # execute the request now
        deferred = getPage(url, timeout=timeout )
        self.deferreds. append(deferred )

        return deferred

        def _callback(self) :
        if len(self.reques ts) self.SIZE:
        queue = self.requests[:self.SIZE]
        self.requests = self.requests[self.SIZE:]
        else:
        queue = self.requests[:]
        self.requests = []

        # execute the requests
        for (url, timeout, deferredHelper) in queue:
        deferred = getPage(url, timeout=timeout )
        self.deferreds. append(deferred )

        deferred.chainD eferred(deferre dHelper)




        Regards Manlio Perillo

        Comment

        • Manlio Perillo

          #5
          Re: Making HTTP requests using Twisted

          Manlio Perillo ha scritto:
          [...]
          Here is a quick example, ABSOLUTELY NOT TESTED:
          >
          class DownloadQueue(o bject):
          SIZE = 50
          >
          def init(self):
          self.requests = [] # queued requests
          self.deferreds = [] # waiting requests
          >
          def addRequest(self , url, timeout):
          if len(self.deferr eds) >= sels.SIZE:
          # wait for completion of all previous requests
          DeferredList(se lf.deferreds
          ).addCallback(s elf._callback)
          self.deferreds = []
          The deferreds list should be cleared in the _callback method, not here.
          Please note that probably there are other bugs.


          Regards Manlio Perillo

          Comment

          Working...