wchar_t

Collapse
This topic is closed.
X
X
 
  • Time
  • Show
Clear All
new posts
  • James Brown

    #1

    wchar_t

    could someone please tell me when the wchar_t type was introduced into
    the C language (and with what version).....pe rhaps it was introduced
    as an extension by alot of compiler venders before it became official?

    I am also interested in finding out what first prompted the introduction of
    this type -
    was it Unicode or did wchar_t happen before Unicode came into existence?

    thanks,
    James


  • Emmanuel Delahaye

    #2
    Re: wchar_t

    James Brown a écrit :[color=blue]
    > could someone please tell me when the wchar_t type was introduced into
    > the C language (and with what version)[/color]

    Addendum 1995

    --
    A+

    Emmanuel Delahaye

    Comment

    • those who know me have no need of my name

      #3
      Re: wchar_t

      in comp.lang.c i read:
      [color=blue]
      >could someone please tell me when the wchar_t type was introduced into
      >the C language (and with what version)[/color]

      in the original standard, in 1989. though it was less than useful until
      amd1 was adopted in 1995, and some might say remains less than successful.

      --
      a signature

      Comment

      • lawrence.jones@ugs.com

        #4
        Re: wchar_t

        James Brown <dont_bother> wrote:[color=blue]
        >
        > I am also interested in finding out what first prompted the introduction of
        > this type -
        > was it Unicode or did wchar_t happen before Unicode came into existence?[/color]

        It was large character sets in general. At the time, the prevalent
        large character sets and encodings were IBM's DBCS (the double-byte
        version of EBCDIC), JIS 208, JIS 212, ISO 2022, SJIS, and EUC. Work had
        begun on what would become ISO 10646, but it was caught up in political
        and technical turmoil between those who insisted that 32 bits were
        required and those who thought that 16 were more than enough and far
        more efficient. The latter camp had just broken away and started work
        on a competing standard, Unicode. (Fortunately for everyone, cooler
        heads prevailed and ISO 10646 and Unicode were eventually harmonized to
        the point that most people now think they're the same thing.)

        -Larry Jones

        Let's pretend I already feel terrible about it, and that you
        don't need to rub it in any more. -- Calvin

        Comment

        • P.J. Plauger

          #5
          Re: wchar_t

          <lawrence.jones @ugs.com> wrote in message
          news:e8t153-vrb.ln1@jones.h omeip.net...
          [color=blue]
          > James Brown <dont_bother> wrote:[color=green]
          >>
          >> I am also interested in finding out what first prompted the introduction
          >> of
          >> this type -
          >> was it Unicode or did wchar_t happen before Unicode came into existence?[/color]
          >
          > It was large character sets in general. At the time, the prevalent
          > large character sets and encodings were IBM's DBCS (the double-byte
          > version of EBCDIC), JIS 208, JIS 212, ISO 2022, SJIS, and EUC. Work had
          > begun on what would become ISO 10646, but it was caught up in political
          > and technical turmoil between those who insisted that 32 bits were
          > required and those who thought that 16 were more than enough and far
          > more efficient. The latter camp had just broken away and started work
          > on a competing standard, Unicode. (Fortunately for everyone, cooler
          > heads prevailed and ISO 10646 and Unicode were eventually harmonized to
          > the point that most people now think they're the same thing.)[/color]

          Right. And the people who thought that 16 bits were more than enough
          and far more efficient are *now* convinced that 21 bits are more
          than enough and far more efficient. I give 'em five years, tops.

          P.J. Plauger
          Dinkumware, Ltd.



          Comment

          • websnarf@gmail.com

            #6
            Re: wchar_t

            P.J. Plauger wrote:[color=blue]
            > <lawrence.jones @ugs.com> wrote in message[color=green]
            > > James Brown <dont_bother> wrote:[color=darkred]
            > >> I am also interested in finding out what first prompted the introduction
            > >> of this type - was it Unicode or did wchar_t happen before Unicode came into
            > >> existence?[/color]
            > >
            > > It was large character sets in general. At the time, the prevalent
            > > large character sets and encodings were IBM's DBCS (the double-byte
            > > version of EBCDIC), JIS 208, JIS 212, ISO 2022, SJIS, and EUC. Work had
            > > begun on what would become ISO 10646, but it was caught up in political
            > > and technical turmoil between those who insisted that 32 bits were
            > > required and those who thought that 16 were more than enough and far
            > > more efficient. The latter camp had just broken away and started work
            > > on a competing standard, Unicode. (Fortunately for everyone, cooler
            > > heads prevailed and ISO 10646 and Unicode were eventually harmonized to
            > > the point that most people now think they're the same thing.)[/color]
            >
            > Right. And the people who thought that 16 bits were more than enough
            > and far more efficient are *now* convinced that 21 bits are more
            > than enough and far more efficient. I give 'em five years, tops.[/color]

            I'll take the other side of any bet you care to form based on that
            statement. (Certainly I'll be recording this message for the
            archives.)

            BTW, the people who thought 32 bits (actually 31bits) was that way to
            go *also* agree that (almost) 21 bits are more than enough. Currently
            less than 17 bits are being used today (dominated by the east asian
            characters), and the growth rate appears to be not worse than a
            thousand new characters added per year. The kinds of things they are
            considering these days are invented character sets (like an
            accessibility alphabet called "blissymbolics" , the script used for
            Klingon and Elvish in the Lord of the Rings series ...) and really
            obscure historical symbols (apparently "old hungarian" used an alphabet
            that doesn't survive to today except by obscure historians in Hungary.)
            This seems to be very asymptotic to me.

            The problem with the 16 bits people (i.e., Microsoft, Sun, and I think
            IBM) is that they were so stupid as to think that Asia wouldn't really
            need complete character sets. They were basically being passively
            racist. But, of course, money talks and there is a lot of commerce in
            and with east asia, so they had to be accomodating. It turns out that
            17 bits appears to be the right answer to get them all, but leaving no
            expandability left at all is clearly insane. Having 21 bits means they
            literally have more than 15 times as much space left over than what
            they are currently using (again, remembering, that they're already
            covering the really "big" east asian character sets).

            The only remaining controversy (that I can tell) is the aliasing of
            characters between the three major east asian languages. From what I'm
            told, people in those countries don't seem to care about the subtle
            problems that causes (you can't quote one language within another
            unless you use some meta data, like a font change), and have gone full
            steam ahead with dropping Big5 and adopting Unicode pretty pervasively.

            You think they'll run out in 5 years? Personally, I think they're
            done.

            --
            Paul Hsieh
            Pobox has been discontinued as a separate service, and all existing customers moved to the Fastmail platform.



            Comment

            • P.J. Plauger

              #7
              Re: wchar_t

              <websnarf@gmail .com> wrote in message
              news:1132394800 .251667.28050@g 43g2000cwa.goog legroups.com...
              [color=blue]
              > The only remaining controversy (that I can tell) is the aliasing of
              > characters between the three major east asian languages. From what I'm
              > told, people in those countries don't seem to care about the subtle
              > problems that causes (you can't quote one language within another
              > unless you use some meta data, like a font change), and have gone full
              > steam ahead with dropping Big5 and adopting Unicode pretty pervasively.
              >
              > You think they'll run out in 5 years? Personally, I think they're
              > done.[/color]

              Here's a coarse scale or two, just from personal experience.

              -- Number of address bits required to address a "large" memory:

              1960 15 IBM 7090
              1970 20 IBM 360
              1980 25 VAX 11/780
              1990 30 various
              2000 35 various

              -- Number of bits required to represent a (commonly used)
              character set:

              1960 6 numerous vendor-specific codes
              1970 7 7-bit ASCII
              1980 8 extended ASCII
              1990 16 DBCS and others
              2000 21 Unicode

              I could make a similar table of "barely adequate" communication
              speeds, which also continue to expand exponentially.

              So long as you think in terms of linear increases in demand
              for bytes or characters, it's easy to believe at each stage
              that you're through expanding. After all, you currently have
              a bit of headroom, and what possible need can there be for
              much larger programs/character sets?

              I personally can't imagine that people will ever want to
              define common attribute bits for, say:

              -- roman, italic, bold, underscore
              -- red, green, blue
              -- point size
              -- font

              But if we did, each attribute bit would double the number
              of effective character codes, wouldn't it?

              Nor can I imagine that a large government like China might
              thumb its nose at an international standard and, say,
              require a parallel set of many ISO 10646 codes.

              For over 40 years I've been reading regular articles by
              pundits who explain why larger/faster hardware is a waste
              of time and will never sell. They've all been wrong. And
              the further back in time you look, the greater the redshift
              in the predictions.

              So, you may well be right that the need for larger
              character sets has finally come to an end. I'll wait
              and see. Meanwhile, I make sure that the code I write
              will work with 32- (not 31-) bit character sets. With
              any luck, the code will have adequate capacity until
              I retire...

              P.J. Plauger
              Dinkumware, Ltd.



              Comment

              • James Brown

                #8
                Re: wchar_t

                <lawrence.jones @ugs.com> wrote in message
                news:e8t153-vrb.ln1@jones.h omeip.net...[color=blue]
                > James Brown <dont_bother> wrote:[color=green]
                >>
                >> I am also interested in finding out what first prompted the introduction
                >> of
                >> this type -
                >> was it Unicode or did wchar_t happen before Unicode came into existence?[/color]
                >
                > It was large character sets in general. At the time, the prevalent
                > large character sets and encodings were IBM's DBCS (the double-byte
                > version of EBCDIC), JIS 208, JIS 212, ISO 2022, SJIS, and EUC. Work had
                > begun on what would become ISO 10646, but it was caught up in political
                > and technical turmoil between those who insisted that 32 bits were
                > required and those who thought that 16 were more than enough and far
                > more efficient. The latter camp had just broken away and started work
                > on a competing standard, Unicode. (Fortunately for everyone, cooler
                > heads prevailed and ISO 10646 and Unicode were eventually harmonized to
                > the point that most people now think they're the same thing.)
                >
                > -Larry Jones
                >
                > Let's pretend I already feel terrible about it, and that you
                > don't need to rub it in any more. -- Calvin[/color]

                thanks! (to everyone) for the very informative answers..

                cheers,
                James


                Comment

                • Skarmander

                  #9
                  [OT] Re: wchar_t

                  I'll mark it OT, since we've left C behind quite a bit by now.

                  P.J. Plauger wrote:[color=blue]
                  > <websnarf@gmail .com> wrote in message
                  > news:1132394800 .251667.28050@g 43g2000cwa.goog legroups.com...
                  >
                  >[color=green]
                  >>The only remaining controversy (that I can tell) is the aliasing of
                  >>characters between the three major east asian languages. From what I'm
                  >>told, people in those countries don't seem to care about the subtle
                  >>problems that causes (you can't quote one language within another
                  >>unless you use some meta data, like a font change), and have gone full
                  >>steam ahead with dropping Big5 and adopting Unicode pretty pervasively.
                  >>
                  >>You think they'll run out in 5 years? Personally, I think they're
                  >>done.[/color]
                  >
                  >
                  > Here's a coarse scale or two, just from personal experience.
                  >
                  > -- Number of address bits required to address a "large" memory:
                  >
                  > 1960 15 IBM 7090
                  > 1970 20 IBM 360
                  > 1980 25 VAX 11/780
                  > 1990 30 various
                  > 2000 35 various
                  >[/color]
                  Nice, but this misses a point: there is an upper limit. Address bits
                  will not continue to grow indefinitely, because there is an upper limit
                  to the amount of information that will fit in the universe. Or maybe
                  there isn't, but then we're talking a radical shift in physics, which
                  may happen but doesn't allow for fair comparison anymore.
                  [color=blue]
                  > -- Number of bits required to represent a (commonly used)
                  > character set:
                  >
                  > 1960 6 numerous vendor-specific codes
                  > 1970 7 7-bit ASCII
                  > 1980 8 extended ASCII
                  > 1990 16 DBCS and others
                  > 2000 21 Unicode
                  >
                  > I could make a similar table of "barely adequate" communication
                  > speeds, which also continue to expand exponentially.
                  >[/color]
                  But again: it can't go on forever. The question here, therefore, is
                  whether we've reached the end of the line, not whether exponential
                  expansion is happening.
                  [color=blue]
                  > So long as you think in terms of linear increases in demand
                  > for bytes or characters, it's easy to believe at each stage
                  > that you're through expanding. After all, you currently have
                  > a bit of headroom, and what possible need can there be for
                  > much larger programs/character sets?
                  >[/color]
                  Don't think this question hasn't been asked, unlike those people who
                  asserted that "640K ought to be enough for anybody" (which Bill Gates
                  famously never said) or "16 bits ought to be enough, since it's better
                  than wasting 32 bits". Unicode doesn't say "21 bits ought to be enough
                  for anybody". It can say "21 bits is enough for every character known to
                  man", because it is. Unlike memory, communication speed and a host of
                  other things that keep growing, there is a conceivable upper limit, and
                  it is not that unreasonable to state we're close to it.
                  [color=blue]
                  > I personally can't imagine that people will ever want to
                  > define common attribute bits for, say:
                  >
                  > -- roman, italic, bold, underscore
                  > -- red, green, blue
                  > -- point size
                  > -- font
                  >
                  > But if we did, each attribute bit would double the number
                  > of effective character codes, wouldn't it?
                  >[/color]

                  That's why Unicode doesn't work that way, and no character set ever has.
                  They encode *characters*, not *glyphs*. A glyph is what you see on your
                  screen, and it may have many nice properties by which it is affected,
                  including the formatting characteristics you describe. But a Roman
                  capital letter A is a Roman capital letter A, no matter what style,
                  color, size or font it happens to be displayed in. Being able to leave
                  these things unstated will always remain useful.

                  Actually, "glyph sets" were (and probably still are) in common use for
                  display on dumb terminals with hardwired character sets (and probably
                  some applications for not so dumb terminals, too). Remember when the
                  character set was 7-bit ASCII and the terminals extended this to an
                  8-bit glyph set with the upper bit meaning "reverse video"? That's this.

                  The point is, effective comparison stops being useful at this point,
                  because you've shifted the way you look at what a code point represents.
                  As the Unicode FAQ itself states:

                  "Both Unicode and ISO 10646 have policies in place that formally limit
                  future code assignment to the integer range that can be expressed with
                  current UTF-16 (0 to 1,114,111). Even if other encoding forms (i.e.
                  other UTFs) can represent larger intergers, these policies mean that all
                  encoding forms will always represent the same set of characters. Over a
                  million possible codes is far more than enough for the goal of Unicode
                  of encoding characters, not glyphs. Unicode is not designed to encode
                  arbitrary data. If you wanted, for example, to give each 'instance of a
                  character on paper throughout history' its own code, you might need
                  trillions or quadrillions of such codes; noble as this effort might be,
                  you would not use Unicode for such an encoding."

                  Here's a more interesting thing to think about than adding "blink" bits:
                  suppose we encounter extraterrestria l cultures one day, and we want to
                  synch character sets eventually... *Then* Unicode may become
                  insufficient. But I don't think it would be fair to blame the current
                  standard for that.
                  [color=blue]
                  > Nor can I imagine that a large government like China might
                  > thumb its nose at an international standard and, say,
                  > require a parallel set of many ISO 10646 codes.
                  >[/color]
                  It already thumbs its nose to some extent. Unicode is still viewed with
                  great suspicion in some parts of the Eastern world, and alternate
                  character sets continue to be in use. But the Chinese government can
                  require of ISO 10646 what it wants; it's not likely to get it if it
                  can't be supported by technical requirements, as opposed to politics.
                  Maybe you can slip in one character that's spurious that way, but not a
                  few thousand. Maybe when the Chinese achieve global domination and
                  abolish our preposterous 21-bit standards, but not before.
                  [color=blue]
                  > For over 40 years I've been reading regular articles by
                  > pundits who explain why larger/faster hardware is a waste
                  > of time and will never sell. They've all been wrong. And
                  > the further back in time you look, the greater the redshift
                  > in the predictions.
                  >[/color]
                  These arguments do not cleanly translate to character sets, your little
                  tables notwithstanding . The upper limit may not be 21 bits, but if
                  that's not the upper limit, it's pretty close to it in orders of
                  magnitude. If people one day decide to abandon the concept of "character
                  set" and go crazy stuffing all sorts of attributes in it (adopting
                  "glyph sets"), that's a clear change in application, unlike increased
                  hardware capacity. It will be fueled by the *ability* to use such sets
                  efficiently, not the *need* to do this.
                  [color=blue]
                  > So, you may well be right that the need for larger
                  > character sets has finally come to an end. I'll wait
                  > and see. Meanwhile, I make sure that the code I write
                  > will work with 32- (not 31-) bit character sets. With
                  > any luck, the code will have adequate capacity until
                  > I retire...
                  >[/color]
                  Fortunately for you, writing code that can handle both 21-bit and 32-bit
                  character sets is hardly a challenge, given the current state of
                  computer hardware. Even if Unicode had to grow someday (which would have
                  to mean a new standard, of course), it wouldn't exactly be hard to
                  implement, at least not as far as code point size is concerned.

                  S.

                  Comment

                  • websnarf@gmail.com

                    #10
                    Re: wchar_t

                    P.J. Plauger wrote:[color=blue]
                    > <websnarf@gmail .com> wrote in message[color=green]
                    > > The only remaining controversy (that I can tell) is the aliasing of
                    > > characters between the three major east asian languages. From what I'm
                    > > told, people in those countries don't seem to care about the subtle
                    > > problems that causes (you can't quote one language within another
                    > > unless you use some meta data, like a font change), and have gone full
                    > > steam ahead with dropping Big5 and adopting Unicode pretty pervasively.
                    > >
                    > > You think they'll run out in 5 years? Personally, I think they're
                    > > done.[/color]
                    >
                    > Here's a coarse scale or two, just from personal experience.
                    >
                    > -- Number of address bits required to address a "large" memory:
                    >
                    > 1960 15 IBM 7090
                    > 1970 20 IBM 360
                    > 1980 25 VAX 11/780
                    > 1990 30 various
                    > 2000 35 various
                    >
                    > -- Number of bits required to represent a (commonly used)
                    > character set:
                    >
                    > 1960 6 numerous vendor-specific codes[/color]

                    Used only by computer scientists. (Commerce on computing being
                    non-existent.)
                    [color=blue]
                    > 1970 7 7-bit ASCII[/color]

                    Used only in english speaking countries.
                    [color=blue]
                    > 1980 8 extended ASCII[/color]

                    Used only in english, and *some* european countries.
                    [color=blue]
                    > 1990 16 DBCS and others[/color]

                    A nonsensical hack.
                    [color=blue]
                    > 2000 21 Unicode[/color]

                    Used in 100% of all computer using countries (and built to scale to
                    those that don't).

                    The only potential for future growth here will come from the SETI
                    project.
                    [color=blue]
                    > I could make a similar table of "barely adequate" communication
                    > speeds, which also continue to expand exponentially.
                    >
                    > So long as you think in terms of linear increases in demand
                    > for bytes or characters, it's easy to believe at each stage
                    > that you're through expanding. After all, you currently have
                    > a bit of headroom, and what possible need can there be for
                    > much larger programs/character sets?[/color]

                    There is nowhere to scale, and the head room is overkill. We would
                    have to add at least 16 languages of similar complexity to the
                    east-asian ones before the encoding space was at risk.
                    [color=blue]
                    > I personally can't imagine that people will ever want to
                    > define common attribute bits for, say:
                    >
                    > -- roman, italic, bold, underscore
                    > -- red, green, blue
                    > -- point size
                    > -- font
                    >
                    > But if we did, each attribute bit would double the number
                    > of effective character codes, wouldn't it?[/color]

                    So you haven't read anything about Unicode at all have you? Unicode
                    does *not* specify meta-information. Those kinds of data will never be
                    put into the Unicode standard, and are not considered part of the text
                    data that Unicode specifies.

                    This also belies an ignorance of what Unicode is specifying. Do you
                    think it makes sense to have the accent of one character in a different
                    font or size than its base character? Even if you wanted to encode
                    this (which I think the east-asians may need in some cases of cross
                    multi-language applications), obviously such meta-data specification
                    would be encoding as escaped *modes*. This is easily encoded in the
                    "private data area" ranges in application specific ways. But most
                    people use meta-display formatting languages, like HTML, or the Open
                    document format, or MS Word, something like that to encode such things
                    today.
                    [color=blue]
                    > Nor can I imagine that a large government like China might
                    > thumb its nose at an international standard and, say,
                    > require a parallel set of many ISO 10646 codes.[/color]

                    Why would they do this? The closest thing to China setting policy on
                    anything regarding computing standards is their adoption of Red Flag
                    Linux. Linux uses Unicode as its internationaliz ation mechanism. I
                    don't think China wants to give up on the commerce that relies on this
                    standardization (i.e., all of it.)
                    [color=blue]
                    > For over 40 years I've been reading regular articles by
                    > pundits who explain why larger/faster hardware is a waste
                    > of time and will never sell. They've all been wrong. And
                    > the further back in time you look, the greater the redshift
                    > in the predictions.[/color]

                    That is because they always underestimate the scale and growth in the
                    problem being solved. By analogy you are suggesting that human
                    languages and the character sets we use will be increasing over time in
                    an increasing and exponential way similar to the growth of programming
                    applications.
                    [color=blue]
                    > So, you may well be right that the need for larger
                    > character sets has finally come to an end. I'll wait
                    > and see. Meanwhile, I make sure that the code I write
                    > will work with 32- (not 31-) bit character sets.[/color]

                    Are you going to invent your own standard? UTF-32 encodes 31 bits (the
                    top bit is assumed to be 0, otherwise an encoding error can be
                    assumed). UTF-8 encodes at most 31 bits (this is a physical encoding
                    limitation). And UTF-16 encodes a little under 21 bits (again, a
                    physical encoding limitation). The only *valid* encodings are the
                    intersection of these which is essentially the UTF-16 encoding.
                    [color=blue]
                    > [...] With any luck, the code will have adequate capacity until I retire...[/color]

                    Also, just arbitrarily thinking "characters are 32 bits" are less that
                    useful to people who actually want to encode and use Unicode data. For
                    example, string comparison and collation cannot be done with a simple
                    byte comparison, and character counts do not correspond to the length
                    of the encoded data. If you don't encode actual Unicode semantics
                    (i.e., you use wchar_t instead) then "adequate" is not something anyone
                    is going to consider your implementation.

                    --
                    Paul Hsieh
                    Pobox has been discontinued as a separate service, and all existing customers moved to the Fastmail platform.



                    Comment

                    • Richard Tobin

                      #11
                      Re: wchar_t

                      In article <1132433097.194 380.145260@g49g 2000cwa.googleg roups.com>,
                      <websnarf@gmail .com> wrote:[color=blue]
                      >And UTF-16 encodes a little under 21 bits[/color]

                      A little over 20 bits would be more accurate: 2^20 + 2^16.

                      -- Richard

                      Comment

                      • P.J. Plauger

                        #12
                        Re: wchar_t

                        <websnarf@gmail .com> wrote in message
                        news:1132433097 .194380.145260@ g49g2000cwa.goo glegroups.com.. .
                        [color=blue][color=green]
                        >> But if we did, each attribute bit would double the number
                        >> of effective character codes, wouldn't it?[/color]
                        >
                        > So you haven't read anything about Unicode at all have you?[/color]

                        Actually, I have.
                        [color=blue]
                        > Unicode
                        > does *not* specify meta-information. Those kinds of data will never be
                        > put into the Unicode standard, and are not considered part of the text
                        > data that Unicode specifies.[/color]

                        What, never? You may very well be right.
                        [color=blue]
                        > This also belies an ignorance of what Unicode is specifying. Do you
                        > think it makes sense to have the accent of one character in a different
                        > font or size than its base character?[/color]

                        Does it make sense to have several different ways to express the
                        same "character" , some involving multiple codes in arbitrary order?
                        Particluarly when there's a one-element version that does the job?
                        Who would do a thing like that in an international standard?
                        [color=blue][color=green]
                        >> Nor can I imagine that a large government like China might
                        >> thumb its nose at an international standard and, say,
                        >> require a parallel set of many ISO 10646 codes.[/color]
                        >
                        > Why would they do this? The closest thing to China setting policy on
                        > anything regarding computing standards is their adoption of Red Flag
                        > Linux. Linux uses Unicode as its internationaliz ation mechanism. I
                        > don't think China wants to give up on the commerce that relies on this
                        > standardization (i.e., all of it.)[/color]

                        That's not what I've heard.
                        [color=blue][color=green]
                        >> For over 40 years I've been reading regular articles by
                        >> pundits who explain why larger/faster hardware is a waste
                        >> of time and will never sell. They've all been wrong. And
                        >> the further back in time you look, the greater the redshift
                        >> in the predictions.[/color]
                        >
                        > That is because they always underestimate the scale and growth in the
                        > problem being solved.[/color]

                        Uh huh.
                        [color=blue]
                        > By analogy you are suggesting that human
                        > languages and the character sets we use will be increasing over time in
                        > an increasing and exponential way similar to the growth of programming
                        > applications.[/color]

                        Yep.
                        [color=blue][color=green]
                        >> So, you may well be right that the need for larger
                        >> character sets has finally come to an end. I'll wait
                        >> and see. Meanwhile, I make sure that the code I write
                        >> will work with 32- (not 31-) bit character sets.[/color]
                        >
                        > Are you going to invent your own standard?[/color]

                        No. But I've already invented my own worst-case *machinery*
                        for handling a variety of standards. Differen thing.
                        [color=blue]
                        > UTF-32 encodes 31 bits (the
                        > top bit is assumed to be 0, otherwise an encoding error can be
                        > assumed). UTF-8 encodes at most 31 bits (this is a physical encoding
                        > limitation). And UTF-16 encodes a little under 21 bits (again, a
                        > physical encoding limitation). The only *valid* encodings are the
                        > intersection of these which is essentially the UTF-16 encoding.[/color]

                        At the moment yes. A few years ago, it was UCS-2.
                        [color=blue][color=green]
                        >> [...] With any luck, the code will have adequate capacity until I
                        >> retire...[/color]
                        >
                        > Also, just arbitrarily thinking "characters are 32 bits" are less that
                        > useful to people who actually want to encode and use Unicode data.[/color]

                        May be true. I didn't say I was.

                        P.J. Plauger
                        Dinkumware, Ltd.



                        Comment

                        • P.J. Plauger

                          #13
                          Re: wchar_t

                          "Richard Tobin" <richard@cogsci .ed.ac.uk> wrote in message
                          news:dlo79b$r11 $1@pc-news.cogsci.ed. ac.uk...
                          [color=blue]
                          > In article <1132433097.194 380.145260@g49g 2000cwa.googleg roups.com>,
                          > <websnarf@gmail .com> wrote:[color=green]
                          >>And UTF-16 encodes a little under 21 bits[/color]
                          >
                          > A little over 20 bits would be more accurate: 2^20 + 2^16.[/color]

                          Right. And the minimum number of real bits needed to express
                          20.087463 bits is...?

                          P.J. Plauger
                          Dinkumware, Ltd.



                          Comment

                          • websnarf@gmail.com

                            #14
                            Re: wchar_t

                            Richard Tobin wrote:[color=blue]
                            > In article <1132433097.194 380.145260@g49g 2000cwa.googleg roups.com>,
                            > <websnarf@gmail .com> wrote:[color=green]
                            > >And UTF-16 encodes a little under 21 bits[/color]
                            >
                            > A little over 20 bits would be more accurate: 2^20 + 2^16.[/color]

                            Ah yes, I misremembered this. But you forgot to subtract out the
                            escape hole itself:

                            2^20 + 2^16 - 2*2^(20/2).

                            Then you can take into account that U+FFFF is always illegal, and at
                            that only one of the two encodings: (xFFFE) or (xFEFF) can be legal in
                            any single given stream (what this means is that once decoded, U+FEFF
                            is legal (and a basically content-free code point), while U+FFFE is
                            not):

                            2^20 + 2^16 - 2^11 - 2.

                            All these complications coming from UTF-16, but which have to be
                            adopted by the other encodings (except the 0xFEFF nonsense) just to
                            make them all consistent.

                            Then I don't know how you want count the unassigned code points that
                            are clearly within the range of certain code point categories. You
                            *know* that those values will never be assigned and never have meaning,
                            but they are not explicitely marked as illegal.

                            --
                            Paul Hsieh
                            Pobox has been discontinued as a separate service, and all existing customers moved to the Fastmail platform.



                            Comment

                            • Richard Tobin

                              #15
                              Re: wchar_t

                              In article <8bKdnXt8vOn6Ke LenZ2dnUVZ_sidn Z2d@giganews.co m>,
                              P.J. Plauger <pjp@dinkumware .com> wrote:
                              [color=blue][color=green][color=darkred]
                              >>>And UTF-16 encodes a little under 21 bits[/color]
                              >>
                              >> A little over 20 bits would be more accurate: 2^20 + 2^16.[/color]
                              >
                              >Right. And the minimum number of real bits needed to express
                              >20.087463 bits is...?[/color]

                              What is your point?

                              -- Richard

                              Comment

                              Working...