using mmap on large (> 2 Gig) files

Collapse
This topic is closed.
X
X
 
  • Time
  • Show
Clear All
new posts
  • sturlamolden

    #16
    Re: using mmap on large (> 2 Gig) files


    Martin v. Löwis wrote:
    You know this isn't true in general. It is true for a 32-bit address
    space only.
    Yes, but there are two other aspects:

    1. Many of us use 32-bit architectures. The one who wrote the module
    should have considered why UNIX' mmap and Windows' MapViewOfFile takes
    an offset parameter. As it is now, "mmap.mmap" can be considered
    inadequate on 32 bit architectures.

    2. The OS may be stupid. Mapping a large file may be a major slowdown
    simply because the memory mapping is implemented suboptimally inside
    the OS. For example it may try to load and synchronise huge portions of
    the file that you don't need. This will deplete the amout of free RAM,
    and perhaps result in excessive swapping. "mmap.mmap" is therefore a
    potential "tarpit" on any architecture. Thus, memory mapping more than
    you need is not intelligent, even if you do have a 64 bit processor.
    The missing offset argument is essential for getting adequate
    performance from a memory-mapped file object.

    Comment

    • sturlamolden

      #17
      Re: using mmap on large (> 2 Gig) files


      Donn Cave wrote:
      Wow, you're sure a wizard! Most people would need to look before
      making statements like that.
      I know, but your news-server doesn't honour cancel messages. :)

      Python's mmap does indeed memory map the file into the process image.
      It does not fake memory mapping by means of file seek operations.

      However, "memory mapping" a file by means of fseek() is probably more
      efficient than using UNIX' mmap() or Windows'
      CreateFileMappi ng()/MapViewOfFile() . In Python, we don't always need
      the file memory mapped, we normally just want to use slicing-operators,
      for-loops and other goodies on the file object -- i.e. we just want to
      treat the file as a Python container object. There are many ways of
      achieving that.

      We can implement a container object backed by a binary file just as
      efficient (and possibly even more efficient) without using the OS'
      memory mapping facilities. The major advantage is that we can
      "pseudo-memory map" a lot more than a 32 bit address space can harbour.


      However - as I wrote in another posting - memory-mapping may also be
      used to create shared memory on Windows, and that doesn't fit easily
      into the fseek scheme. But apart from that, I don't see why true memory
      mapping has any real advantage on Python. As long as slicing operators
      work, users will probably not be able to tell the difference.

      There are in any case room for improving Python's mmap object.

      Comment

      • sturlamolden

        #18
        Re: using mmap on large (> 2 Gig) files


        Martin v. Löwis wrote:

        Your news server doesn't honour cancel as well...
        It doesn't need to, why do you think it does?
        This was an extremely stupid question on my side. It needs to be
        flushed after a write because that's how the memory pages mapping the
        file is synchronized with the file. Write ops to the memory mapping
        addresses isn't immediately synchronized with the file on disk. Both
        Windows and UNIX require this. I should think before I write, but I
        realized this after posting and my cancel didn't reach you.
        Read the source, Luke. It uses mmap or MapViewOfFile, depending
        on the platform.
        Yes, indeed.

        Comment

        • Steve Holden

          #19
          Re: using mmap on large (> 2 Gig) files

          sturlamolden wrote:
          [...]
          This was an extremely stupid question on my side.
          I take my hat off to anyone who's prepared to admit this. We all do it,
          but most of us try to ignore the fact.

          regards
          Steve
          --
          Steve Holden +44 150 684 7255 +1 800 494 3119
          Holden Web LLC/Ltd http://www.holdenweb.com
          Skype: holdenweb http://holdenweb.blogspot.com
          Recent Ramblings http://del.icio.us/steve.holden

          Comment

          • Martin v. Löwis

            #20
            Re: using mmap on large (> 2 Gig) files

            sturlamolden schrieb:
            2. The OS may be stupid. Mapping a large file may be a major slowdown
            simply because the memory mapping is implemented suboptimally inside
            the OS. For example it may try to load and synchronise huge portions of
            the file that you don't need.
            Can you give an example of an operating system that behaves that way?
            To my knowledge, all current systems integrating memory mapping somehow
            with the page/buffer caches, using various strategies to write-back
            (or just discard in case of no writes) pages that haven't been used
            for a while.
            The missing offset argument is essential for getting adequate
            performance from a memory-mapped file object.
            I very much question that statement. Do you have any numbers to
            prove it?

            Regards,
            Martin

            Comment

            • Tim Roberts

              #21
              Re: using mmap on large (> 2 Gig) files

              "sturlamold en" <sturlamolden@y ahoo.nowrote:
              >
              >However, "memory mapping" a file by means of fseek() is probably more
              >efficient than using UNIX' mmap() or Windows'
              >CreateFileMapp ing()/MapViewOfFile() .
              My goodness, do I disagree with that! At least on Windows, I/O on a file
              mapped with MapViewOfFile uses the virtual memory pager -- the same
              mechanism used by the swap file. Because it is so heavily used, that is
              some of the most well-optimized code in the system.
              >We can implement a container object backed by a binary file just as
              >efficient (and possibly even more efficient) without using the OS'
              >memory mapping facilities. The major advantage is that we can
              >"pseudo-memory map" a lot more than a 32 bit address space can harbour.
              Both the Unix mmap and the Win32 MapViewOfFile allow a starting byte
              offset. It wouldn't be rocket science to extend Python's mmap to allow
              that.
              >There are in any case room for improving Python's mmap object.
              Here we agree.
              --
              Tim Roberts, timr@probo.com
              Providenza & Boekelheide, Inc.

              Comment

              • Paul Rubin

                #22
                Re: using mmap on large (&gt; 2 Gig) files

                "sturlamold en" <sturlamolden@y ahoo.nowrites:
                However, "memory mapping" a file by means of fseek() is probably more
                efficient than using UNIX' mmap() or Windows'
                CreateFileMappi ng()/MapViewOfFile() .
                Why on would you think that?! It is counterintuitiv e. fseek beyond
                whatever is buffered in stdio (usually no more than 1kbyte or so)
                requires a system call, while mmap is just a memory access.
                In Python, we don't always need the file memory mapped, we normally
                just want to use slicing-operators, for-loops and other goodies on
                the file object -- i.e. we just want to treat the file as a Python
                container object. There are many ways of achieving that.
                Some of the time we want to share the region with other processes.
                Sometimes we just want random access to a big file on disk without
                having to do a lot of context switches seeking around in the file.
                There are in any case room for improving Python's mmap object.
                IMO it should have some kind of IPC locking mechanism added, in
                addition to the offset stuff suggested.

                Comment

                • Chetan

                  #23
                  Re: using mmap on large (&gt; 2 Gig) files

                  Paul Rubin <http://phr.cx@NOSPAM.i nvalidwrites:
                  "sturlamold en" <sturlamolden@y ahoo.nowrites:
                  >However, "memory mapping" a file by means of fseek() is probably more
                  >efficient than using UNIX' mmap() or Windows'
                  >CreateFileMapp ing()/MapViewOfFile() .
                  >
                  Why on would you think that?! It is counterintuitiv e. fseek beyond
                  whatever is buffered in stdio (usually no more than 1kbyte or so)
                  requires a system call, while mmap is just a memory access.
                  And the buffer copy required with every I/O from/to the application.
                  >In Python, we don't always need the file memory mapped, we normally
                  >just want to use slicing-operators, for-loops and other goodies on
                  >the file object -- i.e. we just want to treat the file as a Python
                  >container object. There are many ways of achieving that.
                  >
                  Some of the time we want to share the region with other processes.
                  Sometimes we just want random access to a big file on disk without
                  having to do a lot of context switches seeking around in the file.
                  >
                  >There are in any case room for improving Python's mmap object.
                  >
                  IMO it should have some kind of IPC locking mechanism added, in
                  addition to the offset stuff suggested.
                  The type of IPC required differs depending on who is using the shared region -
                  either another python process or another external program. Apart from the
                  spinlock primitives, other types of synchronization mechanisms are provided by
                  the OS. However, I do see value in providing a shared memory based spinlock
                  mechanism. These services can be built on top of the shared memory
                  infrastructure. I am not sure what kind or real world python applications use
                  it.

                  -Chetan

                  Comment

                  • Paul Rubin

                    #24
                    Re: using mmap on large (&gt; 2 Gig) files

                    Chetan <pandyacus.xspa m@xspam.sbcglob al.netwrites:
                    Why on would you think that?! It is counterintuitiv e. fseek beyond
                    whatever is buffered in stdio (usually no more than 1kbyte or so)
                    requires a system call, while mmap is just a memory access.
                    And the buffer copy required with every I/O from/to the application.
                    Even that can probably be avoided since the mmap region has to start
                    on a page boundary, but anyway regular I/O definitely has to copy the
                    data. For mmap, I'm thinking mostly of the case where the entire file
                    is paged in through most of the program's execution though. That
                    obviously wouldn't apply to every application.
                    IMO it should have some kind of IPC locking mechanism added, in
                    addition to the offset stuff suggested.
                    The type of IPC required differs depending on who is using the
                    shared region - either another python process or another external
                    program. Apart from the spinlock primitives, other types of
                    synchronization mechanisms are provided by the OS. However, I do see
                    value in providing a shared memory based spinlock mechanism.
                    I mean just have an interface to OS locks (Linux futex and whatever
                    the Windows counterpart is) and maybe also a utility function to do a
                    compare-and-swap in user space.

                    Comment

                    • Chetan

                      #25
                      Re: using mmap on large (&gt; 2 Gig) files

                      Paul Rubin <http://phr.cx@NOSPAM.i nvalidwrites:
                      I mean just have an interface to OS locks (Linux futex and whatever
                      the Windows counterpart is) and maybe also a utility function to do a
                      compare-and-swap in user space.
                      There is code for spinlocks, but it allocates the lockword in the process
                      memory. This can be used for thread synchronization , but not for IPC with
                      external python or non-python processes.
                      I found a PyIPC IPC package that seems to provide interface to Sys V shared
                      memory and semaphore - but I just found it, so cannot comment on it at this
                      time.

                      Comment

                      • nnorwitz@gmail.com

                        #26
                        Re: using mmap on large (&gt; 2 Gig) files

                        Martin v. Löwis wrote:
                        sturlamolden schrieb:
                        >
                        And why doesn't Python's mmap take an offset argument to handle large
                        files?
                        >
                        I don't know exactly; the most likely reason is that nobody has
                        contributed code to make it support that. That's, in turn, probably
                        because nobody had the problem yet, or nobody of those who did
                        cared enough to implement and contribute a patch.
                        Or because no one cared enough to test a patch that was produced 2.5
                        years ago (not directed at Martin, just pointing out why the patch
                        stalled).



                        With just a little community support, this can go in. I suppose now
                        that we have the buildbots, we can check in untested code and test it
                        that way. The patch should be reviewed.

                        n

                        Comment

                        • Chetan

                          #27
                          Re: using mmap on large (&gt; 2 Gig) files

                          "nnorwitz@gmail .com" <nnorwitz@gmail .comwrites:
                          Martin v. Löwis wrote:
                          sturlamolden schrieb:
                          And why doesn't Python's mmap take an offset argument to handle large
                          files?
                          I don't know exactly; the most likely reason is that nobody has
                          contributed code to make it support that. That's, in turn, probably
                          because nobody had the problem yet, or nobody of those who did
                          cared enough to implement and contribute a patch.
                          >
                          Or because no one cared enough to test a patch that was produced 2.5
                          years ago (not directed at Martin, just pointing out why the patch
                          stalled).
                          >

                          >
                          With just a little community support, this can go in. I suppose now
                          that we have the buildbots, we can check in untested code and test it
                          that way. The patch should be reviewed.
                          >
                          n
                          I made the changes before I saw this. However, the patch seems to be quite
                          dated and some of the changes are very interesting, especially if they were
                          tested for the special conditions they are supposed to handle and
                          if they were made after some discussion.
                          I can submit my patch as it is, but I am working on making some of the other
                          changes I had in mind for the mmap to be useful.
                          Some of the other changes would make more sense for py3k, if it supports a byte
                          array object, but I haven't looked at py3k at all.

                          Chetan

                          Comment

                          Working...