Help with Tokenizer Please

Collapse
This topic is closed.
X
X
 
  • Time
  • Show
Clear All
new posts
  • electrixnow

    #1

    Help with Tokenizer Please

    I use a code snippet called tokenize that works quite well. But I need
    to modify it so it returns an empty string if two commas are found
    together: ie. string test="1,2,,4,5"

    I would like to have

    tokens[0] = "1"
    tokens[1] = "2"
    tokens[3] = ""
    tokens[4] = "4"
    tokens[5] = "5"

    This would allow me to find missing tokens. The following is the code I
    use.
    Can anyone show me the modification I would need?
    -----------------

    void Tokenize(const string& str,
    vector<string>& tokens,
    const string& delimiters = " ,\t\n")
    {
    string::size_ty pe lastPos = str.find_first_ not_of(delimite rs, 0);
    // Skip delimiters at beginning.
    string::size_ty pe pos = str.find_first_ of(delimiters, lastPos);
    // Find first "non-delimiter".
    while (string::npos != pos || string::npos != lastPos)
    {
    tokens.push_bac k(str.substr(la stPos, pos - lastPos));
    // Found a token, add it to the vector.
    lastPos = str.find_first_ not_of(delimite rs, pos);
    // Skip delimiters. Note the "not_of"
    pos = str.find_first_ of(delimiters, lastPos);
    // Find next "non-delimiter"
    }
    }

    --------------

    Thanks in advance,
    Grant

  • Victor Bazarov

    #2
    Re: Help with Tokenizer Please

    electrixnow wrote:[color=blue]
    > I use a code snippet called tokenize that works quite well. But I need
    > to modify it so it returns an empty string if two commas are found
    > together: ie. string test="1,2,,4,5"
    >
    > I would like to have
    >
    > tokens[0] = "1"
    > tokens[1] = "2"
    > tokens[3] = ""
    > tokens[4] = "4"
    > tokens[5] = "5"
    >
    > This would allow me to find missing tokens. The following is the code
    > I use.
    > Can anyone show me the modification I would need?
    > -----------------
    >
    > void Tokenize(const string& str,
    > vector<string>& tokens,
    > const string& delimiters = " ,\t\n")
    > {
    > string::size_ty pe lastPos = str.find_first_ not_of(delimite rs, 0);
    > // Skip delimiters at beginning.
    > string::size_ty pe pos = str.find_first_ of(delimiters, lastPos);
    > // Find first "non-delimiter".
    > while (string::npos != pos || string::npos != lastPos)
    > {
    > tokens.push_bac k(str.substr(la stPos, pos - lastPos));
    > // Found a token, add it to the vector.
    > lastPos = str.find_first_ not_of(delimite rs, pos);
    > // Skip delimiters. Note the "not_of"[/color]

    I think you shouldn't skip delimiters here. You just need to make
    'lastPos' to be the same as 'pos'. Replace the statement above with

    lastPos = pos;
    [color=blue]
    > pos = str.find_first_ of(delimiters, lastPos);
    > // Find next "non-delimiter"
    > }
    > }
    >
    > --------------[/color]

    I didn't check it. It's just a hunch.

    V
    --
    Please remove capital As from my address when replying by mail


    Comment

    Working...