<uchar.h> provides char16_t and char32_t with fixed, portable UTF-16 and UTF-32 encodings (C11). Conversion functions translate between multibyte (UTF-8) and these encodings, suitable for portable Unicode handling.
Use this instead of wchar_t for Unicode code that must work identically across platforms.
This example converts a UTF-8 string to UTF-32 code points and displays them.
// compile: gcc -o ucharexample ucharexample.c // run: ./ucharexample // description: convert UTF-8 to UTF-32 and iterate code points #include <uchar.h> #include <stdio.h> #include <string.h> #include <locale.h> int main() { setlocale(LC_ALL, ""); const char* utf8 = "héllo"; mbstate_t state = {0}; char32_t c32; size_t n; while (*utf8) { n = mbrtoc32(&c32, utf8, strlen(utf8), &state); if (n == 0 || n == (size_t)-1 || n == (size_t)-2) break; printf("U+%04X (%zu bytes)\n", (unsigned)c32, n); utf8 += n; } return 0; }
Character types:
char16_t: 16-bit Unicode character (UTF-16 code unit)char32_t: 32-bit Unicode character (UTF-32 code point)mbstate_t: multibyte conversion stateUTF-8 to UTF-32 conversion:
mbrtoc32(&c32, s, n, &state): convert next multibyte sequence to UTF-32c32rtomb(s, c32, &state): convert UTF-32 code point to multibyteUTF-8 to UTF-16 conversion:
mbrtoc16(&c16, s, n, &state): convert next multibyte sequence to UTF-16c16rtomb(s, c16, &state): convert UTF-16 code unit to multibyteInitialization and reset:
mbstate_t state = {0}: initialize conversion stateReturn values (from mbrtoc32/mbrtoc16):
(size_t)0: null character converted(size_t)-1: error (invalid sequence)(size_t)-2: incomplete sequence (need more bytes)