creating/modifying sparse files on linux

Collapse
This topic is closed.
X
X
 
  • Time
  • Show
Clear All
new posts
  • draghuram@gmail.com

    #1

    creating/modifying sparse files on linux


    Hi,

    Is there any special support for sparse file handling in python? My
    initial search didn't bring up much (not a thorough search). I wrote
    the following pice of code:

    options.size = 6442450944
    options.ranges = ["4096,1024","30 000,314572800"]
    fd = open("testfile" , "w")
    fd.seek(options .size-1)
    fd.write("a")
    for drange in options.ranges:
    off = int(drange.spli t(",")[0])
    len = int(drange.spli t(",")[1])
    print "off =", off, " len =", len
    fd.seek(off)
    for x in range(len):
    fd.write("a")

    fd.close()

    This piece of code takes very long time and in fact I had to kill it as
    the linux system started doing lot of swapping. Am I doing something
    wrong here? Is there a better way to create/modify sparse files?

    Thanks,
    Raghu.

  • Trent Mick

    #2
    Re: creating/modifying sparse files on linux

    [draghuram@gmail .com wrote][color=blue]
    >
    > Hi,
    >
    > Is there any special support for sparse file handling in python? My
    > initial search didn't bring up much (not a thorough search). I wrote
    > the following pice of code:
    >
    > options.size = 6442450944
    > options.ranges = ["4096,1024","30 000,314572800"]
    > fd = open("testfile" , "w")
    > fd.seek(options .size-1)
    > fd.write("a")
    > for drange in options.ranges:
    > off = int(drange.spli t(",")[0])
    > len = int(drange.spli t(",")[1])
    > print "off =", off, " len =", len
    > fd.seek(off)
    > for x in range(len):
    > fd.write("a")
    >
    > fd.close()
    >
    > This piece of code takes very long time and in fact I had to kill it as
    > the linux system started doing lot of swapping. Am I doing something
    > wrong here? Is there a better way to create/modify sparse files?[/color]

    test_largefile. py in the Python test suite does this kind of thing and
    doesn't take very long for me to run on Linux (SuSE 9.0 box).

    Trent

    --
    Trent Mick
    TrentM@ActiveSt ate.com

    Comment

    • Marc 'BlackJack' Rintsch

      #3
      Re: creating/modifying sparse files on linux

      In <1124304819.090 356.200280@f14g 2000cwb.googleg roups.com>,
      draghuram@gmail .com wrote:
      [color=blue]
      > options.size = 6442450944
      > options.ranges = ["4096,1024","30 000,314572800"]
      > fd = open("testfile" , "w")
      > fd.seek(options .size-1)
      > fd.write("a")
      > for drange in options.ranges:
      > off = int(drange.spli t(",")[0])
      > len = int(drange.spli t(",")[1])
      > print "off =", off, " len =", len
      > fd.seek(off)
      > for x in range(len):
      > fd.write("a")
      >
      > fd.close()
      >
      > This piece of code takes very long time and in fact I had to kill it as
      > the linux system started doing lot of swapping. Am I doing something
      > wrong here? Is there a better way to create/modify sparse files?[/color]

      `range(len)` creates a list of size `len` *in memory* so you are trying to
      build a list with 314,572,800 numbers. That seems to eat up all your RAM
      and causes the swapping.

      You can use `xrange(len)` instead which uses a constant amount of memory.
      But be prepared to wait some time because now you are writing 314,572,800
      characters *one by one* into the file. It would be faster to write larger
      strings in each step.

      Ciao,
      Marc 'BlackJack' Rintsch

      Comment

      • Terry Reedy

        #4
        Re: creating/modifying sparse files on linux


        <draghuram@gmai l.com> wrote in message
        news:1124304819 .090356.200280@ f14g2000cwb.goo glegroups.com.. .[color=blue]
        > Is there any special support for sparse file handling in python?[/color]

        Since I have not heard of such in several years, I suspect not. CPython,
        normally compiled, uses the standard C stdio lib. If your system+C has a
        sparseIO lib, you would probably have to compile specially to use it.
        [color=blue]
        > options.size = 6442450944
        > options.ranges = ["4096,1024","30 000,314572800"][/color]

        options.ranges = [(4096,1024),(30 000,314572800)] # makes below nicer
        [color=blue]
        > fd = open("testfile" , "w")
        > fd.seek(options .size-1)
        > fd.write("a")
        > for drange in options.ranges:
        > off = int(drange.spli t(",")[0])
        > len = int(drange.spli t(",")[1])[/color]

        off,len = map(int, drange.split(", ")) # or
        off,len = [int(s) for s in drange.split(", ")] # or for tuples as suggested
        above
        off,len = drange
        [color=blue]
        > print "off =", off, " len =", len
        > fd.seek(off)
        > for x in range(len):[/color]

        If I read the above right, the 2nd len is 300,000,000+ making the space
        needed for the range list a few gigabytes. I suspect this is where you
        started thrashing ;-). Instead:

        for x in xrange(len): # this is what xrange is for ;-)
        [color=blue]
        > fd.write("a")[/color]

        Without indent, this is syntax error, so if your code ran at all, this
        cannot be an exact copy. Even with xrange fix, 300,000,000 writes will be
        slow. I would expect that an real application should create or accumulate
        chunks larger than single chars.
        [color=blue]
        > fd.close()
        >
        > This piece of code takes very long time and in fact I had to kill it as
        > the linux system started doing lot of swapping. Am I doing something
        > wrong here?[/color]

        See above
        [color=blue]
        > Is there a better way to create/modify sparse files?[/color]

        Unless you can access builting facilities, create your own mapping index.

        Terry J. Reedy



        Comment

        • draghuram@gmail.com

          #5
          Re: creating/modifying sparse files on linux


          Thanks for the info on xrange. Writing single char is just to get going
          quickly. I knew that I would have to improve on that. I would like to
          write chunks of 1MB which would require that I have 1MB string to
          write. Is there any simple way of generating this 1MB string (other
          than keep appending to a string until it reaches 1MB len)? I don't care
          about the actual value of the string itself.

          Thanks,
          Raghu.

          Comment

          • Terry Reedy

            #6
            Re: creating/modifying sparse files on linux


            <draghuram@gmai l.com> wrote in message
            news:1124314877 .646737.266290@ z14g2000cwz.goo glegroups.com.. .[color=blue]
            >
            > Thanks for the info on xrange. Writing single char is just to get going
            > quickly. I knew that I would have to improve on that. I would like to
            > write chunks of 1MB which would require that I have 1MB string to
            > write. Is there any simple way of generating this 1MB string[/color]

            megastring = 1000000*'a' # t < 1 sec on my machine
            [color=blue]
            >(other than keep appending to a string until it reaches 1MB len)?[/color]

            You mean like (unexecuted)
            s = ''
            for i in xrange(1000000) : s += 'a' #?

            This will allocate, copy, and deallocate 1000000 successively longer
            temporary strings and is a noticeable O(n**2) operation. Since strings are
            immutable, you cannot 'append' to them the way you can to lists.

            Terry J. Reedy



            Comment

            • François Pinard

              #7
              Re: creating/modifying sparse files on linux

              [draghuram@gmail .com]
              [color=blue]
              > Is there any simple way of generating this 1MB string (other than keep
              > appending to a string until it reaches 1MB len)?[/color]

              You might of course use 'x' * 1000000 for fairly quickly generating a
              single string holding one million `x'.

              Yet, your idea of generating a sparse file is interesting. I never
              tried it with Python, but would not see why Python would not allow
              it. Did someone ever played with sparse files in Python? (One problem
              with sparse files is that it is next to impossible for a normal user to
              create an exact copy. There is no fast way to read read them either.)

              --
              François Pinard http://pinard.progiciels-bpi.ca

              Comment

              • Bengt Richter

                #8
                Re: creating/modifying sparse files on linux

                On 17 Aug 2005 11:53:39 -0700, "draghuram@gmai l.com" <draghuram@gmai l.com> wrote:
                [color=blue]
                >
                >Hi,
                >
                >Is there any special support for sparse file handling in python? My
                >initial search didn't bring up much (not a thorough search). I wrote
                >the following pice of code:
                >
                >options.size = 6442450944
                >options.rang es = ["4096,1024","30 000,314572800"]
                >fd = open("testfile" , "w")
                >fd.seek(option s.size-1)
                >fd.write("a" )
                >for drange in options.ranges:
                > off = int(drange.spli t(",")[0])
                > len = int(drange.spli t(",")[1])
                > print "off =", off, " len =", len
                > fd.seek(off)
                > for x in range(len):
                > fd.write("a")
                >
                >fd.close()
                >
                >This piece of code takes very long time and in fact I had to kill it as
                >the linux system started doing lot of swapping. Am I doing something
                >wrong here? Is there a better way to create/modify sparse files?
                >
                >Thanks[/color]
                I'm unclear as to what your goal is. Do you just need an object that provides
                an interface like a file object, but internally is more efficient than an
                a normal file object when you access it as above[1], or do you need to create
                a real file and record all the bytes in full (with what default for gaps?)
                on disk, so that it can be opened by another program and read as an ordinary file?

                Some operating system file systems may have some support for virtual zero-block runs
                and lazy allocation/representation of non-zero blocks in files. It's easy to imagine
                the rudiments, but I don't know of such a file system, not having looked ;-)

                You could write your own "sparse-file"-representation object, and maybe use pickle
                for persistence. Or maybe you could use zipfiles. The kind of data you are creating above
                would probably compress really well ;-)

                [1] writing 314+ million identical bytes one by one is silly, of course ;-)
                BTW, len is a built-in function, and using built-in names for variables
                is frowned upon as a bug-prone practice.

                Regards,
                Bengt Richter

                Comment

                • Benji York

                  #9
                  Re: creating/modifying sparse files on linux

                  Terry Reedy wrote:[color=blue]
                  > megastring = 1000000*'a' # t < 1 sec on my machine[/color]
                  [color=blue][color=green]
                  >>(other than keep appending to a string until it reaches 1MB len)?[/color]
                  >
                  > You mean like (unexecuted)
                  > s = ''
                  > for i in xrange(1000000) : s += 'a' #?
                  >
                  > This will allocate, copy, and deallocate 1000000 successively longer
                  > temporary strings and is a noticeable O(n**2) operation.[/color]

                  Not exactly. CPython 2.4 added an optimization of "+=" for strings.
                  The for loop above takes about 1 second do execute on my machine. You
                  are correct in that it will take *much* longer on 2.3.
                  --
                  Benji York

                  Comment

                  • draghuram@gmail.com

                    #10
                    Re: creating/modifying sparse files on linux


                    My goal is very simple. Have a mechanism to create sparse files and
                    modify them by writing arbitratry ranges of bytes at arbitrary offsets.
                    I did get the information I want (xrange instead of range, and a simple
                    way to generate 1Mb string in memory). Thanks for pointing out about
                    using "len" as variable. It is indeed silly.

                    My only assumption from underlying OS/file system is that if I seek
                    past end of file and write some data, it doesn't generate blocks for
                    data in between. This is indeed true on Linux (I tested on ext3).

                    Thanks,
                    Raghu.

                    Comment

                    • Mike Meyer

                      #11
                      Re: creating/modifying sparse files on linux

                      "draghuram@gmai l.com" <draghuram@gmai l.com> writes:
                      [color=blue]
                      > My goal is very simple. Have a mechanism to create sparse files and
                      > modify them by writing arbitratry ranges of bytes at arbitrary offsets.
                      > I did get the information I want (xrange instead of range, and a simple
                      > way to generate 1Mb string in memory). Thanks for pointing out about
                      > using "len" as variable. It is indeed silly.
                      >
                      > My only assumption from underlying OS/file system is that if I seek
                      > past end of file and write some data, it doesn't generate blocks for
                      > data in between. This is indeed true on Linux (I tested on ext3).[/color]

                      This better be true for anything claiming to be Unix. The results on
                      systems that break this aren't pretty.

                      <mike
                      --
                      Mike Meyer <mwm@mired.or g> http://www.mired.org/home/mwm/
                      Independent WWW/Perforce/FreeBSD/Unix consultant, email for more information.

                      Comment

                      Working...