re beginner

Collapse
This topic is closed.
X
X
 
  • Time
  • Show
Clear All
new posts
  • SuperHik

    #1

    re beginner

    hi all,

    I'm trying to understand regex for the first time, and it would be very
    helpful to get an example. I have an old(er) script with the following
    task - takes a string I copy-pasted and wich always has the same format:
    [color=blue][color=green][color=darkred]
    >>> print stuff[/color][/color][/color]
    Yellow hat 2 Blue shirt 1
    White socks 4 Green pants 1
    Blue bag 4 Nice perfume 3
    Wrist watch 7 Mobile phone 4
    Wireless cord! 2 Building tools 3
    One for the money 7 Two for the show 4
    [color=blue][color=green][color=darkred]
    >>> stuff[/color][/color][/color]
    'Yellow hat\t2\tBlue shirt\t1\nWhite socks\t4\tGreen pants\t1\nBlue
    bag\t4\tNice perfume\t3\nWri st watch\t7\tMobil e phone\t4\nWirel ess
    cord!\t2\tBuild ing tools\t3\nOne for the money\t7\tTwo for the show\t4'

    I want to put items from stuff into a dict like this:[color=blue][color=green][color=darkred]
    >>> print mydict[/color][/color][/color]
    {'Wireless cord!': 2, 'Green pants': 1, 'Blue shirt': 1, 'White socks':
    4, 'Mobile phone': 4, 'Two for the show': 4, 'One for the money': 7,
    'Blue bag': 4, 'Wrist watch': 7, 'Nice perfume': 3, 'Yellow hat': 2,
    'Building tools': 3}

    Here's how I did it:[color=blue][color=green][color=darkred]
    >>> def putindict(items ):[/color][/color][/color]
    .... items = items.replace(' \n', '\t')
    .... items = items.split('\t ')
    .... d = {}
    .... for x in xrange( len(items) ):
    .... if not items[x].isdigit(): d[items[x]] = int(items[x+1])
    .... return d[color=blue][color=green][color=darkred]
    >>>
    >>> mydict = putindict(stuff )[/color][/color][/color]


    I was wondering is there a better way to do it using re module?
    perheps even avoiding this for loop?

    thanks!
  • bearophileHUGS@lycos.com

    #2
    Re: re beginner

    SuperHik wrote:[color=blue]
    > I was wondering is there a better way to do it using re module?
    > perheps even avoiding this for loop?[/color]

    This is a way to do the same thing without REs:

    data = 'Yellow hat\t2\tBlue shirt\t1\nWhite socks\t4\tGreen
    pants\t1\nBlue bag\t4\tNice perfume\t3\nWri st watch\t7\tMobil e
    phone\t4\nWirel ess cord!\t2\tBuild ing tools\t3\nOne for the
    money\t7\tTwo for the show\t4'

    data2 = data.replace("\ n","\t").split( "\t")
    result1 = dict( zip(data2[::2], map(int, data2[1::2])) )

    O if you want to be light:

    from itertools import imap, izip, islice
    data2 = data.replace("\ n","\t").split( "\t")
    strings = islice(data2, 0, len(data), 2)
    numbers = islice(data2, 1, len(data), 2)
    result2 = dict( izip(strings, imap(int, numbers)) )

    Bye,
    bearophile

    Comment

    • faulkner

      #3
      Re: re beginner

      you could write a function which takes a match object and modifies d,
      pass the function to re.sub, and ignore what re.sub returns.

      # untested code
      d = {}
      def record(match):
      s = match.string[match.start() : match.end()]
      i = s.index('\t')
      print s, i # debugging
      d[s[:i]] = int(s[i+1:])
      return ''
      re.sub('\w+\t\d +\t', record, stuff)
      # end code

      it may be a bit faster, but it's very roundabout and difficult to
      debug.

      SuperHik wrote:[color=blue]
      > hi all,
      >
      > I'm trying to understand regex for the first time, and it would be very
      > helpful to get an example. I have an old(er) script with the following
      > task - takes a string I copy-pasted and wich always has the same format:
      >[color=green][color=darkred]
      > >>> print stuff[/color][/color]
      > Yellow hat 2 Blue shirt 1
      > White socks 4 Green pants 1
      > Blue bag 4 Nice perfume 3
      > Wrist watch 7 Mobile phone 4
      > Wireless cord! 2 Building tools 3
      > One for the money 7 Two for the show 4
      >[color=green][color=darkred]
      > >>> stuff[/color][/color]
      > 'Yellow hat\t2\tBlue shirt\t1\nWhite socks\t4\tGreen pants\t1\nBlue
      > bag\t4\tNice perfume\t3\nWri st watch\t7\tMobil e phone\t4\nWirel ess
      > cord!\t2\tBuild ing tools\t3\nOne for the money\t7\tTwo for the show\t4'
      >
      > I want to put items from stuff into a dict like this:[color=green][color=darkred]
      > >>> print mydict[/color][/color]
      > {'Wireless cord!': 2, 'Green pants': 1, 'Blue shirt': 1, 'White socks':
      > 4, 'Mobile phone': 4, 'Two for the show': 4, 'One for the money': 7,
      > 'Blue bag': 4, 'Wrist watch': 7, 'Nice perfume': 3, 'Yellow hat': 2,
      > 'Building tools': 3}
      >
      > Here's how I did it:[color=green][color=darkred]
      > >>> def putindict(items ):[/color][/color]
      > ... items = items.replace(' \n', '\t')
      > ... items = items.split('\t ')
      > ... d = {}
      > ... for x in xrange( len(items) ):
      > ... if not items[x].isdigit(): d[items[x]] = int(items[x+1])
      > ... return d[color=green][color=darkred]
      > >>>
      > >>> mydict = putindict(stuff )[/color][/color]
      >
      >
      > I was wondering is there a better way to do it using re module?
      > perheps even avoiding this for loop?
      >
      > thanks![/color]

      Comment

      • bearophileHUGS@lycos.com

        #4
        Re: re beginner

        > strings = islice(data2, 0, len(data), 2)[color=blue]
        > numbers = islice(data2, 1, len(data), 2)[/color]

        This probably has to be:

        strings = islice(data2, 0, len(data2), 2)
        numbers = islice(data2, 1, len(data2), 2)

        Sorry,
        bearophile

        Comment

        • Bruno Desthuilliers

          #5
          Re: re beginner

          SuperHik a écrit :[color=blue]
          > hi all,
          >
          > I'm trying to understand regex for the first time, and it would be very
          > helpful to get an example. I have an old(er) script with the following
          > task - takes a string I copy-pasted and wich always has the same format:
          >[color=green][color=darkred]
          > >>> print stuff[/color][/color]
          > Yellow hat 2 Blue shirt 1
          > White socks 4 Green pants 1
          > Blue bag 4 Nice perfume 3
          > Wrist watch 7 Mobile phone 4
          > Wireless cord! 2 Building tools 3
          > One for the money 7 Two for the show 4
          >[color=green][color=darkred]
          > >>> stuff[/color][/color]
          > 'Yellow hat\t2\tBlue shirt\t1\nWhite socks\t4\tGreen pants\t1\nBlue
          > bag\t4\tNice perfume\t3\nWri st watch\t7\tMobil e phone\t4\nWirel ess
          > cord!\t2\tBuild ing tools\t3\nOne for the money\t7\tTwo for the show\t4'
          >
          > I want to put items from stuff into a dict like this:[color=green][color=darkred]
          > >>> print mydict[/color][/color]
          > {'Wireless cord!': 2, 'Green pants': 1, 'Blue shirt': 1, 'White socks':
          > 4, 'Mobile phone': 4, 'Two for the show': 4, 'One for the money': 7,
          > 'Blue bag': 4, 'Wrist watch': 7, 'Nice perfume': 3, 'Yellow hat': 2,
          > 'Building tools': 3}
          >
          > Here's how I did it:[color=green][color=darkred]
          > >>> def putindict(items ):[/color][/color]
          > ... items = items.replace(' \n', '\t')
          > ... items = items.split('\t ')
          > ... d = {}
          > ... for x in xrange( len(items) ):
          > ... if not items[x].isdigit(): d[items[x]] = int(items[x+1])
          > ... return d[color=green][color=darkred]
          > >>>
          > >>> mydict = putindict(stuff )[/color][/color]
          >
          >
          > I was wondering is there a better way to do it using re module?
          > perheps even avoiding this for loop?[/color]

          There are better ways. One of them avoids the for loop, and even the re
          module:

          def to_dict(items):
          items = items.replace(' \t', '\n').split('\n ')
          return dict(zip(items[::2], map(int, items[1::2])))

          HTH

          Comment

          • Bruno Desthuilliers

            #6
            Re: re beginner

            bearophileHUGS@ lycos.com a écrit :[color=blue][color=green]
            >>strings = islice(data2, 0, len(data), 2)
            >>numbers = islice(data2, 1, len(data), 2)[/color]
            >
            >
            > This probably has to be:
            >
            > strings = islice(data2, 0, len(data2), 2)
            > numbers = islice(data2, 1, len(data2), 2)[/color]

            try with islice(data2, 0, None, 2)

            Comment

            • John Machin

              #7
              Re: re beginner

              On 5/06/2006 10:38 AM, Bruno Desthuilliers wrote:[color=blue]
              > SuperHik a écrit :[color=green]
              >> hi all,
              >>
              >> I'm trying to understand regex for the first time, and it would be
              >> very helpful to get an example. I have an old(er) script with the
              >> following task - takes a string I copy-pasted and wich always has the
              >> same format:
              >>[color=darkred]
              >> >>> print stuff[/color]
              >> Yellow hat 2 Blue shirt 1
              >> White socks 4 Green pants 1
              >> Blue bag 4 Nice perfume 3
              >> Wrist watch 7 Mobile phone 4
              >> Wireless cord! 2 Building tools 3
              >> One for the money 7 Two for the show 4
              >>[color=darkred]
              >> >>> stuff[/color]
              >> 'Yellow hat\t2\tBlue shirt\t1\nWhite socks\t4\tGreen pants\t1\nBlue
              >> bag\t4\tNice perfume\t3\nWri st watch\t7\tMobil e phone\t4\nWirel ess
              >> cord!\t2\tBuild ing tools\t3\nOne for the money\t7\tTwo for the show\t4'
              >>
              >> I want to put items from stuff into a dict like this:[color=darkred]
              >> >>> print mydict[/color]
              >> {'Wireless cord!': 2, 'Green pants': 1, 'Blue shirt': 1, 'White
              >> socks': 4, 'Mobile phone': 4, 'Two for the show': 4, 'One for the
              >> money': 7, 'Blue bag': 4, 'Wrist watch': 7, 'Nice perfume': 3, 'Yellow
              >> hat': 2, 'Building tools': 3}
              >>
              >> Here's how I did it:[color=darkred]
              >> >>> def putindict(items ):[/color]
              >> ... items = items.replace(' \n', '\t')
              >> ... items = items.split('\t ')
              >> ... d = {}
              >> ... for x in xrange( len(items) ):
              >> ... if not items[x].isdigit(): d[items[x]] = int(items[x+1])
              >> ... return d[color=darkred]
              >> >>>
              >> >>> mydict = putindict(stuff )[/color]
              >>
              >>
              >> I was wondering is there a better way to do it using re module?
              >> perheps even avoiding this for loop?[/color]
              >
              > There are better ways. One of them avoids the for loop, and even the re
              > module:
              >
              > def to_dict(items):
              > items = items.replace(' \t', '\n').split('\n ')[/color]

              In case there are leading/trailing spaces on the keys:

              items = [x.strip() for x in items.replace(' \t', '\n').split('\n ')]
              [color=blue]
              > return dict(zip(items[::2], map(int, items[1::2])))
              >
              > HTH[/color]

              Fantastic -- at least for the OP's carefully copied-and-pasted input.
              Meanwhile back in the real world, there might be problems with multiple
              tabs used for 'prettiness' instead of 1 tab, non-integer values, etc etc.
              In that case a loop approach that validated as it went and was able to
              report the position and contents of any invalid input might be better.

              Comment

              • Paul McGuire

                #8
                Re: re beginner

                "John Machin" <sjmachin@lexic on.net> wrote in message
                news:4483665A.2 06@lexicon.net. ..[color=blue]
                > Fantastic -- at least for the OP's carefully copied-and-pasted input.
                > Meanwhile back in the real world, there might be problems with multiple
                > tabs used for 'prettiness' instead of 1 tab, non-integer values, etc etc.
                > In that case a loop approach that validated as it went and was able to
                > report the position and contents of any invalid input might be better.[/color]

                Yeah, for that you'd need more like a real parser... hey, wait a minute!
                What about pyparsing?!

                Here's a pyparsing version. The definition of the parsing patterns takes
                little more than the re definition does - the bulk of the rest of the code
                is parsing/scanning the input and reporting the results.

                The pyparsing home page is at http://pyparsing.wikispaces.com.

                -- Paul


                stuff = 'Yellow hat\t2\tBlue shirt\t1\nWhite socks\t4\tGreen pants\t1\nBlue
                bag\t4\tNice perfume\t3\nWri st watch\t7\tMobil e phone\t4\nWirel ess
                cord!\t2\tBuild ing tools\t3\nOne for the money\t7\tTwo for the show\t4'
                print "Original input string:"
                print stuff
                print

                from pyparsing import *

                # define low-level elements for parsing
                itemWord = Word(alphas, alphanums+".!?" )
                itemDesc = OneOrMore(itemW ord)
                integer = Word(nums)

                # add parse action to itemDesc to merge separate words into single string
                itemDesc.setPar seAction( lambda s,l,t: " ".join(t) )

                # define macro element for an entry
                entry = itemDesc.setRes ultsName("item" ) + integer.setResu ltsName("qty")

                # scan through input string for entry's, print out their named fields
                print "Results when scanning for entries:"
                for t,s,e in entry.scanStrin g(stuff):
                print t.item,t.qty
                print

                # parse entire string, building ParseResults with dict-like access
                results = dictOf( itemDesc, integer ).parseString(s tuff)
                print "Results when parsing entries as a dict:"
                print "Keys:", results.keys()
                for item in results.items() :
                print item
                for k in results.keys():
                print k,"=", results[k]


                prints:

                Original input string:
                Yellow hat 2 Blue shirt 1
                White socks 4 Green pants 1
                Blue bag 4 Nice perfume 3
                Wrist watch 7 Mobile phone 4
                Wireless cord! 2 Building tools 3
                One for the money 7 Two for the show 4

                Results when scanning for entries:
                Yellow hat 2
                Blue shirt 1
                White socks 4
                Green pants 1
                Blue bag 4
                Nice perfume 3
                Wrist watch 7
                Mobile phone 4
                Wireless cord! 2
                Building tools 3
                One for the money 7
                Two for the show 4

                Results when parsing entries as a dict:
                Keys: ['Wireless cord!', 'Green pants', 'Blue shirt', 'White socks', 'Mobile
                phone', 'Two for the show', 'One for the money', 'Blue bag', 'Wrist watch',
                'Nice perfume', 'Yellow hat', 'Building tools']
                ('Wireless cord!', '2')
                ('Green pants', '1')
                ('Blue shirt', '1')
                ('White socks', '4')
                ('Mobile phone', '4')
                ('Two for the show', '4')
                ('One for the money', '7')
                ('Blue bag', '4')
                ('Wrist watch', '7')
                ('Nice perfume', '3')
                ('Yellow hat', '2')
                ('Building tools', '3')
                Wireless cord! = 2
                Green pants = 1
                Blue shirt = 1
                White socks = 4
                Mobile phone = 4
                Two for the show = 4
                One for the money = 7
                Blue bag = 4
                Wrist watch = 7
                Nice perfume = 3
                Yellow hat = 2
                Building tools = 3


                Comment

                • John Machin

                  #9
                  Re: re beginner

                  On 5/06/2006 10:07 AM, Paul McGuire wrote:[color=blue]
                  > "John Machin" <sjmachin@lexic on.net> wrote in message
                  > news:4483665A.2 06@lexicon.net. ..[color=green]
                  >> Fantastic -- at least for the OP's carefully copied-and-pasted input.
                  >> Meanwhile back in the real world, there might be problems with multiple
                  >> tabs used for 'prettiness' instead of 1 tab, non-integer values, etc etc.
                  >> In that case a loop approach that validated as it went and was able to
                  >> report the position and contents of any invalid input might be better.[/color]
                  >
                  > Yeah, for that you'd need more like a real parser... hey, wait a minute!
                  > What about pyparsing?!
                  >
                  > Here's a pyparsing version. The definition of the parsing patterns takes
                  > little more than the re definition does - the bulk of the rest of the code
                  > is parsing/scanning the input and reporting the results.
                  >[/color]

                  [big snip]

                  I didn't see any evidence of error handling in there anywhere.


                  Comment

                  • Paul McGuire

                    #10
                    Re: re beginner

                    "John Machin" <sjmachin@lexic on.net> wrote in message
                    news:448399AF.9 030001@lexicon. net...[color=blue]
                    > On 5/06/2006 10:07 AM, Paul McGuire wrote:[color=green]
                    > > "John Machin" <sjmachin@lexic on.net> wrote in message
                    > > news:4483665A.2 06@lexicon.net. ..[color=darkred]
                    > >> Fantastic -- at least for the OP's carefully copied-and-pasted input.
                    > >> Meanwhile back in the real world, there might be problems with multiple
                    > >> tabs used for 'prettiness' instead of 1 tab, non-integer values, etc[/color][/color][/color]
                    etc.[color=blue][color=green][color=darkred]
                    > >> In that case a loop approach that validated as it went and was able to
                    > >> report the position and contents of any invalid input might be better.[/color]
                    > >
                    > > Yeah, for that you'd need more like a real parser... hey, wait a minute!
                    > > What about pyparsing?!
                    > >
                    > > Here's a pyparsing version. The definition of the parsing patterns[/color][/color]
                    takes[color=blue][color=green]
                    > > little more than the re definition does - the bulk of the rest of the[/color][/color]
                    code[color=blue][color=green]
                    > > is parsing/scanning the input and reporting the results.
                    > >[/color]
                    >
                    > [big snip]
                    >
                    > I didn't see any evidence of error handling in there anywhere.
                    >
                    >[/color]
                    Pyparsing has a certain amount of error reporting built in, raising a
                    ParseException when a mismatch occurs.

                    This particular "grammar" is actually pretty error-tolerant. To force an
                    error, I replaced "One for the money" with "1 for the money", and here is
                    the exception reported by pyparsing, along with a diagnostic method,
                    markInputline:


                    stuff = 'Yellow hat\t2\tBlue shirt\t1\nWhite socks\t4\tGreen pants\t1\nBlue
                    bag\t4\tNice perfume\t3\nWri st watch\t7\tMobil e phone\t4\nWirel ess
                    cord!\t2\tBuild ing tools\t3\nOne for the money\t7\tTwo for the show\t4'
                    badstuff = 'Yellow hat\t2\tBlue shirt\t1\nWhite socks\t4\tGreen
                    pants\t1\nBlue bag\t4\tNice perfume\t3\nWri st watch\t7\tMobil e
                    phone\t4\nWirel ess cord!\t2\tBuild ing tools\t3\n1 for the money\t7\tTwo for
                    the show\t4'
                    pattern = dictOf( itemDesc, integer ) + stringEnd
                    print pattern.parseSt ring(stuff)
                    print
                    try:
                    print pattern.parseSt ring(badstuff)
                    except ParseException, pe:
                    print pe
                    print pe.markInputlin e()

                    Gives:
                    [['Yellow hat', '2'], ['Blue shirt', '1'], ['White socks', '4'], ['Green
                    pants', '1'], ['Blue bag', '4'], ['Nice perfume', '3'], ['Wrist watch',
                    '7'], ['Mobile phone', '4'], ['Wireless cord!', '2'], ['Building tools',
                    '3'], ['One for the money', '7'], ['Two for the show', '4']]

                    Expected stringEnd (at char 210), (line:6, col:1)[color=blue]
                    >!<1 for the money 7 Two for the show 4[/color]

                    -- Paul


                    Comment

                    • Bruno Desthuilliers

                      #11
                      Re: re beginner

                      John Machin a écrit :[color=blue]
                      > On 5/06/2006 10:38 AM, Bruno Desthuilliers wrote:
                      >[color=green]
                      >> SuperHik a écrit :
                      >>[color=darkred]
                      >>> hi all,
                      >>>[/color][/color][/color]
                      (snip)
                      [color=blue][color=green][color=darkred]
                      >>> I have an old(er) script with the
                      >>> following task - takes a string I copy-pasted and wich always has the
                      >>> same format:
                      >>>[/color][/color][/color]
                      (snip)[color=blue][color=green][color=darkred]
                      >>>[/color]
                      >> def to_dict(items):
                      >> items = items.replace(' \t', '\n').split('\n ')[/color]
                      >
                      >
                      > In case there are leading/trailing spaces on the keys:[/color]

                      There aren't. Test passes.

                      (snip)
                      [color=blue]
                      > Fantastic -- at least for the OP's carefully copied-and-pasted input.[/color]

                      That was the spec, and my code passes the test.
                      [color=blue]
                      > Meanwhile back in the real world,[/color]

                      The "real world" is mostly defined by customer's test set (is that the
                      correct translation for "jeu d'essai" ?). Code passes the test. period.
                      [color=blue]
                      > there might be problems with multiple
                      > tabs used for 'prettiness' instead of 1 tab, non-integer values, etc etc.[/color]

                      Which means that the spec and the customer's test set is wrong. Not my
                      responsability. Any way, I refuse to change anything in the parsing
                      algorithm before having another test set.
                      [color=blue]
                      > In that case a loop approach that validated as it went and was able to
                      > report the position and contents of any invalid input might be better.[/color]

                      One doesn't know what *will* be better without actual facts. You can be
                      right (and, from my experience, you probably are !-), *but* you can be
                      wrong as well. Until you have a correct spec and test data set on which
                      the code fails, writing any other code is a waste of time. Better to
                      work on other parts of the system, and come back on this if and when the
                      need arise.

                      <ot>
                      Kind of reminds me of a former employer that paid me 2 full monthes to
                      work on a very hairy data migration script (the original data set was so
                      f... up and incoherent even a human parser could barely make any sens of
                      it), before discovering than none of the users of the old system was
                      interested in migrating that part of the data. Talk about a waste of
                      time and money...
                      </ot>

                      Now FWIW, there's actually something else bugging me with this code : it
                      loads the whole data set in memory. It's ok for a few lines, but
                      obviously wrong if one is to parse huge files. *That* would be the first
                      thing I would change - it takes a couple of minutes to do so no real
                      waste of time, but it obviously imply rethinking the API, which is
                      better done yet than when client code will have been written.

                      My 2 cents....

                      Comment

                      • Fredrik Lundh

                        #12
                        Re: re beginner

                        John Machin wrote:
                        [color=blue]
                        > Fantastic -- at least for the OP's carefully copied-and-pasted input.
                        > Meanwhile back in the real world, there might be problems with multiple
                        > tabs used for 'prettiness' instead of 1 tab, non-integer values, etc etc.[/color]

                        yeah, that's probably why the OP stated "which always has the same format".

                        and the "trying to understand regex for the first time, and it would be
                        very helpful to get an example" part was obviously mostly irrelevant to
                        the "smarter than thou" crowd; only one thread contributor was silly
                        enough to actually provide an RE-based example.

                        </F>

                        Comment

                        • Fredrik Lundh

                          #13
                          Re: re beginner

                          SuperHik wrote:
                          [color=blue]
                          > I'm trying to understand regex for the first time, and it would be very
                          > helpful to get an example. I have an old(er) script with the following
                          > task - takes a string I copy-pasted and wich always has the same format:
                          >[color=green][color=darkred]
                          > >>> print stuff[/color][/color]
                          > Yellow hat 2 Blue shirt 1
                          > White socks 4 Green pants 1
                          > Blue bag 4 Nice perfume 3
                          > Wrist watch 7 Mobile phone 4
                          > Wireless cord! 2 Building tools 3
                          > One for the money 7 Two for the show 4
                          >[color=green][color=darkred]
                          > >>> stuff[/color][/color]
                          > 'Yellow hat\t2\tBlue shirt\t1\nWhite socks\t4\tGreen pants\t1\nBlue
                          > bag\t4\tNice perfume\t3\nWri st watch\t7\tMobil e phone\t4\nWirel ess
                          > cord!\t2\tBuild ing tools\t3\nOne for the money\t7\tTwo for the show\t4'[/color]

                          the first thing you need to do is to figure out exactly what the syntax
                          is. given your example, the format of the items you are looking for
                          seems to be "some text" followed by a tab character followed by an integer.

                          a initial attempt would be "\w+\t\d+" (one or more word characters,
                          followed by a tab, followed by one or more digits). to try this out,
                          you can do:
                          [color=blue][color=green][color=darkred]
                          >>> re.findall('\w+ \t\d+', stuff)[/color][/color][/color]
                          ['hat\t2', 'shirt\t1', 'socks\t4', ...]

                          as you can see, using \w+ isn't good enough here; the "keys" in this
                          case may contain whitespace as well, and findall simply skips stuff that
                          doesn't match the pattern. if we assume that a key consists of words
                          and spaces, we can replace the single \w with [\w ] (either word
                          character or space), and get
                          [color=blue][color=green][color=darkred]
                          >>> re.findall('[\w ]+\t\d+', stuff)[/color][/color][/color]
                          ['Yellow hat\t2', 'Blue shirt\t1', 'White socks\t4', ...]

                          which looks a bit better. however, if you check the output carefully,
                          you'll notice that the "Wireless cord!" entry is missing: the "!" isn't
                          a letter or a digit. the easiest way to fix this is to look for
                          "non-tab characters" instead, using "[^\t]" (this matches anything
                          except a tab):
                          [color=blue][color=green][color=darkred]
                          >>> len(re.findall( '[\w ]+\t\d+', stuff))[/color][/color][/color]
                          11[color=blue][color=green][color=darkred]
                          >>> len(re.findall( '[^\t]+\t\d+', stuff))[/color][/color][/color]
                          12

                          now, to turn this into a dictionary, you could split the returned
                          strings on a tab character (\t), but RE provides a better mechanism:
                          capturing groups. by adding () to the pattern string, you can mark the
                          sections you want returned:
                          [color=blue][color=green][color=darkred]
                          >>> re.findall('([^\t]+)\t(\d+)', stuff)[/color][/color][/color]
                          [('Yellow hat', '2'), ('Blue shirt', '1'), ('White socks', ...]

                          turning this into a dictionary is trivial:
                          [color=blue][color=green][color=darkred]
                          >>> dict(re.findall ('([^\t]+)\t(\d+)', stuff))[/color][/color][/color]
                          {'Green pants': '1', 'Blue shirt': '1', 'White socks': ...}[color=blue][color=green][color=darkred]
                          >>> len(dict(re.fin dall('([^\t]+)\t(\d+)', stuff)))[/color][/color][/color]
                          12

                          or, in function terms:

                          def putindict(items ):
                          return dict(re.findall ('([^\t]+)\t(\d+)', stuff))

                          hope this helps!

                          </F>

                          Comment

                          • Bruno Desthuilliers

                            #14
                            Re: re beginner

                            Fredrik Lundh a écrit :[color=blue]
                            > John Machin wrote:
                            >[color=green]
                            >> Fantastic -- at least for the OP's carefully copied-and-pasted input.
                            >> Meanwhile back in the real world, there might be problems with
                            >> multiple tabs used for 'prettiness' instead of 1 tab, non-integer
                            >> values, etc etc.[/color]
                            >
                            >
                            > yeah, that's probably why the OP stated "which always has the same format".[/color]

                            Lol.
                            [color=blue]
                            > and the "trying to understand regex for the first time, and it would be
                            > very helpful to get an example" part[/color]

                            Yeps, I missed that part when answering yesterday. My bad.

                            Comment

                            • John Machin

                              #15
                              Re: re beginner

                              On 5/06/2006 7:47 PM, Fredrik Lundh wrote:[color=blue]
                              > John Machin wrote:
                              >[color=green]
                              >> Fantastic -- at least for the OP's carefully copied-and-pasted input.
                              >> Meanwhile back in the real world, there might be problems with
                              >> multiple tabs used for 'prettiness' instead of 1 tab, non-integer
                              >> values, etc etc.[/color]
                              >
                              > yeah, that's probably why the OP stated "which always has the same format".
                              >[/color]

                              Such statements by users are in the the same category as "The cheque is
                              in the mail" and "Of course I'll still love you in the morning".

                              Comment

                              Working...