Site Tools


h-uchar

<uchar.h>

<uchar.h> provides char16_t and char32_t with fixed, portable UTF-16 and UTF-32 encodings (C11). Conversion functions translate between multibyte (UTF-8) and these encodings, suitable for portable Unicode handling.

Use this instead of wchar_t for Unicode code that must work identically across platforms.

Example

This example converts a UTF-8 string to UTF-32 code points and displays them.

// compile: gcc -o ucharexample ucharexample.c
// run: ./ucharexample
// description: convert UTF-8 to UTF-32 and iterate code points
 
#include <uchar.h>
#include <stdio.h>
#include <string.h>
#include <locale.h>
 
int main() {
    setlocale(LC_ALL, "");
 
    const char* utf8 = "héllo";
    mbstate_t state = {0};
    char32_t c32;
    size_t n;
 
    while (*utf8) {
        n = mbrtoc32(&c32, utf8, strlen(utf8), &state);
        if (n == 0 || n == (size_t)-1 || n == (size_t)-2) break;
        printf("U+%04X (%zu bytes)\n", (unsigned)c32, n);
        utf8 += n;
    }
 
    return 0;
}

Common types and functions

Character types:

  • char16_t: 16-bit Unicode character (UTF-16 code unit)
  • char32_t: 32-bit Unicode character (UTF-32 code point)
  • mbstate_t: multibyte conversion state

UTF-8 to UTF-32 conversion:

  • mbrtoc32(&c32, s, n, &state): convert next multibyte sequence to UTF-32
  • c32rtomb(s, c32, &state): convert UTF-32 code point to multibyte

UTF-8 to UTF-16 conversion:

  • mbrtoc16(&c16, s, n, &state): convert next multibyte sequence to UTF-16
  • c16rtomb(s, c16, &state): convert UTF-16 code unit to multibyte

Initialization and reset:

  • mbstate_t state = {0}: initialize conversion state
  • Re-initialize state before each new string to process multi-byte sequences

Return values (from mbrtoc32/mbrtoc16):

  • (size_t)0: null character converted
  • (size_t)-1: error (invalid sequence)
  • (size_t)-2: incomplete sequence (need more bytes)
  • Other positive value: bytes consumed from input
h-uchar.md · Last modified: by 127.0.0.1