UTF-8到Unicode代码点 [英] UTF-8 to Unicode Code Points
问题描述
是否存在将UTF-8更改为Unicode并将非特殊字符保留为普通字母和数字的功能?
Is there a function that will change UTF-8 to Unicode leaving non special characters as normal letters and numbers?
即德语单词tchüß"将被翻译为"tch \ 20AC \ 21AC"(请注意,我正在编写Unicode代码).
ie the German word "tchüß" would be rendered as something like "tch\20AC\21AC" (please note that I am making the Unicode codes up).
我正在尝试以下功能,但是尽管此功能在ASCII 32-127上运行良好,但对于双字节字符似乎失败了:
I am experimenting with the following function, but although this one works well with ASCII 32-127, it seems to fail for double byte chars:
function strToHex ($string)
{
$hex = '';
for ($i = 0; $i < mb_strlen ($string, "utf-8"); $i++)
{
$id = ord (mb_substr ($string, $i, 1, "utf-8"));
$hex .= ($id <= 128) ? mb_substr ($string, $i, 1, "utf-8") : "&#" . $id . ";";
}
return ($hex);
}
有什么想法吗?
找到解决方案:PHP ord()函数不适用于双字节字符.改用: http://nl.php.net/manual/zh/function.ord.php#78032
EDIT 2: Found solution: The PHP ord() function does not work for double byte chars. Use instead: http://nl.php.net/manual/en/function.ord.php#78032
推荐答案
使用iconv可以将一个字符集转换为另一个字符集:
Converting one character set to another can be done with iconv:
http://php.net/manual/en/function.iconv.php
请注意,UTF已经是Unicode编码.
Note that UTF is already an Unicode encoding.
另一种方法是简单地使用具有正确字符集的htmlentities:
Another way is simply using htmlentities with the right character set:
http://php.net/manual/en/function.htmlentities.php
这篇关于UTF-8到Unicode代码点的文章就介绍到这了,希望我们推荐的答案对大家有所帮助,也希望大家多多支持IT屋!