forked from boostorg/regex
1127 lines
47 KiB
HTML
1127 lines
47 KiB
HTML
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.0 Transitional//EN">
|
|
<html>
|
|
<head>
|
|
<title>Boost.Regex: Localisation</title>
|
|
<meta http-equiv="Content-Type" content="text/html; charset=iso-8859-1">
|
|
<link rel="stylesheet" type="text/css" href="../../../boost.css">
|
|
</head>
|
|
<body>
|
|
<P>
|
|
<TABLE id="Table1" cellSpacing="1" cellPadding="1" width="100%" border="0">
|
|
<TR>
|
|
<td valign="top" width="300">
|
|
<h3><a href="../../../index.htm"><img height="86" width="277" alt="C++ Boost" src="../../../c++boost.gif" border="0"></a></h3>
|
|
</td>
|
|
<TD width="353">
|
|
<H1 align="center">Boost.Regex</H1>
|
|
<H2 align="center">Localisation</H2>
|
|
</TD>
|
|
<td width="50">
|
|
<h3><a href="index.html"><img height="45" width="43" alt="Boost.Regex Index" src="uarrow.gif" border="0"></a></h3>
|
|
</td>
|
|
</TR>
|
|
</TABLE>
|
|
</P>
|
|
<HR>
|
|
<p></p>
|
|
<P>Boost.regex provides extensive support for run-time localization, the
|
|
localization model used can be split into two parts: front-end and back-end.</P>
|
|
<P></P>
|
|
<P>Front-end localization deals with everything which the user sees - error
|
|
messages, and the regular expression syntax itself. For example a French
|
|
application could change [[:word:]] to [[:mot:]] and \w to \m. Modifying the
|
|
front end locale requires active support from the developer, by providing the
|
|
library with a message catalogue to load, containing the localized strings.
|
|
Front-end locale is affected by the LC_MESSAGES category only.
|
|
</P>
|
|
<P>Back-end localization deals with everything that occurs after the expression
|
|
has been parsed - in other words everything that the user does not see or
|
|
interact with directly. It deals with case conversion, collation, and character
|
|
class membership. The back-end locale does not require any intervention from
|
|
the developer - the library will acquire all the information it requires for
|
|
the current locale from the underlying operating system / run time library.
|
|
This means that if the program user does not interact with regular expressions
|
|
directly - for example if the expressions are embedded in your C++ code - then
|
|
no explicit localization is required, as the library will take care of
|
|
everything for you. For example embedding the expression [[:word:]]+ in your
|
|
code will always match a whole word, if the program is run on a machine with,
|
|
for example, a Greek locale, then it will still match a whole word, but in
|
|
Greek characters rather than Latin ones. The back-end locale is affected by the
|
|
LC_TYPE and LC_COLLATE categories.
|
|
</P>
|
|
<P>There are three separate localization mechanisms supported by boost.regex:</P>
|
|
<P><I>Win32 localization model.</I>
|
|
</P>
|
|
<P>This is the default model when the library is compiled under Win32, and is
|
|
encapsulated by the traits class w32_regex_traits. When this model is in effect
|
|
there is a single global locale as defined by the user's control panel
|
|
settings, and returned by GetUserDefaultLCID. All the settings used by
|
|
boost.regex are acquired directly from the operating system bypassing the C run
|
|
time library. Front-end localization requires a resource dll, containing a
|
|
string table with the user-defined strings. The traits class exports the
|
|
function:
|
|
</P>
|
|
<P>static std::string set_message_catalogue(const std::string& s);
|
|
</P>
|
|
<P>which needs to be called with a string identifying the name of the resource
|
|
dll, <I>before</I> your code compiles any regular expressions (but not
|
|
necessarily before you construct any <I>reg_expression</I> instances):
|
|
</P>
|
|
<P>boost::w32_regex_traits<char>::set_message_catalogue("mydll.dll");
|
|
</P>
|
|
<P>Note that this API sets the dll name for <I>both</I> the narrow and wide
|
|
character specializations of w32_regex_traits.
|
|
</P>
|
|
<P>This model does not currently support thread specific locales (via
|
|
SetThreadLocale under Windows NT), the library provides full Unicode support
|
|
under NT, under Windows 9x the library degrades gracefully - characters 0 to
|
|
255 are supported, the remainder are treated as "unknown" graphic characters.
|
|
</P>
|
|
<P><I>C localization model.</I>
|
|
</P>
|
|
<P>This is the default model when the library is compiled under an operating
|
|
system other than Win32, and is encapsulated by the traits class <I>c_regex_traits</I>,
|
|
Win32 users can force this model to take effect by defining the pre-processor
|
|
symbol BOOST_REGEX_USE_C_LOCALE. When this model is in effect there is a single
|
|
global locale, as set by <I>setlocale</I>. All settings are acquired from your
|
|
run time library, consequently Unicode support is dependent upon your run time
|
|
library implementation. Front end localization requires a POSIX message
|
|
catalogue. The traits class exports the function:
|
|
</P>
|
|
<P>static std::string set_message_catalogue(const std::string& s);
|
|
</P>
|
|
<P>which needs to be called with a string identifying the name of the message
|
|
catalogue, <I>before</I> your code compiles any regular expressions (but not
|
|
necessarily before you construct any <I>reg_expression</I> instances):
|
|
</P>
|
|
<P>boost::c_regex_traits<char>::set_message_catalogue("mycatalogue");
|
|
</P>
|
|
<P>Note that this API sets the dll name for <I>both</I> the narrow and wide
|
|
character specializations of c_regex_traits. If your run time library does not
|
|
support POSIX message catalogues, then you can either provide your own
|
|
implementation of <nl_types.h> or define BOOST_RE_NO_CAT to disable
|
|
front-end localization via message catalogues.
|
|
</P>
|
|
<P>Note that calling <I>setlocale</I> invalidates all compiled regular
|
|
expressions, calling <TT>setlocale(LC_ALL, "C")</TT> will make this library
|
|
behave equivalent to most traditional regular expression libraries including
|
|
version 1 of this library.
|
|
</P>
|
|
<P><I><TT>C++ </TT></I><I>localization</I><I><TT> </TT></I><I>model</I><I><TT>.</TT></I>
|
|
</P>
|
|
<P>This model is only in effect if the library is built with the pre-processor
|
|
symbol BOOST_REGEX_USE_CPP_LOCALE defined. When this model is in effect each
|
|
instance of reg_expression<> has its own instance of std::locale, class
|
|
reg_expression<> also has a member function <I>imbue</I> which allows the
|
|
locale for the expression to be set on a per-instance basis. Front end
|
|
localization requires a POSIX message catalogue, which will be loaded via the
|
|
std::messages facet of the expression's locale, the traits class exports the
|
|
symbol:
|
|
</P>
|
|
<P>static std::string set_message_catalogue(const std::string& s);
|
|
</P>
|
|
<P>which needs to be called with a string identifying the name of the message
|
|
catalogue, <I>before</I> your code compiles any regular expressions (but not
|
|
necessarily before you construct any <I>reg_expression</I> instances):
|
|
</P>
|
|
<P>boost::cpp_regex_traits<char>::set_message_catalogue("mycatalogue");
|
|
</P>
|
|
<P>Note that calling reg_expression<>::imbue will invalidate any expression
|
|
currently compiled in that instance of reg_expression<>. This model is
|
|
the one which closest fits the ethos of the C++ standard library, however it is
|
|
the model which will produce the slowest code, and which is the least well
|
|
supported by current standard library implementations, for example I have yet
|
|
to find an implementation of std::locale which supports either message
|
|
catalogues, or locales other than "C" or "POSIX".
|
|
</P>
|
|
<P>Finally note that if you build the library with a non-default localization
|
|
model, then the appropriate pre-processor symbol (BOOST_REGEX_USE_C_LOCALE or
|
|
BOOST_REGEX_USE_CPP_LOCALE) must be defined both when you build the support
|
|
library, and when you include <boost/regex.hpp> or
|
|
<boost/cregex.hpp> in your code. The best way to ensure this is to add
|
|
the #define to <boost/regex/user.hpp>.
|
|
</P>
|
|
<P><I>Providing a message catalogue:</I>
|
|
</P>
|
|
<P>In order to localize the front end of the library, you need to provide the
|
|
library with the appropriate message strings contained either in a resource
|
|
dll's string table (Win32 model), or a POSIX message catalogue (C or C++
|
|
models). In the latter case the messages must appear in message set zero of the
|
|
catalogue. The messages and their id's are as follows:
|
|
<BR>
|
|
|
|
</P>
|
|
<P>
|
|
<TABLE id="Table2" cellSpacing="0" cellPadding="6" width="624" border="0">
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">Message id
|
|
</TD>
|
|
<TD vAlign="top" width="32%">Meaning
|
|
</TD>
|
|
<TD vAlign="top" width="29%">Default value
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">101
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character used to start a sub-expression.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"("
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">102
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character used to end a sub-expression
|
|
declaration.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">")"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">103
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character used to denote an end of line
|
|
assertion.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"$"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">104
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character used to denote the start of line
|
|
assertion.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"^"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">105
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character used to denote the "match any character
|
|
expression".
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"."
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">106
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The match zero or more times repetition operator.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"*"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">107
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The match one or more repetition operator.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"+"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">108
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The match zero or one repetition operator.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"?"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">109
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character set opening character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"["
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">110
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character set closing character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"]"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">111
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The alternation operator.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"|"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">112
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The escape character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"\\"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">113
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The hash character (not currently used).
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"#"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">114
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The range operator.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"-"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">115
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The repetition operator opening character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"{"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">116
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The repetition operator closing character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"}"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">117
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The digit characters.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"0123456789"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">118
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the word boundary assertion.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"b"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">119
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the non-word boundary assertion.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"B"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">120
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the word-start boundary assertion.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"<"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">121
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the word-end boundary assertion.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">">"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">122
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents any word character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"w"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">123
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents a non-word character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"W"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">124
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents a start of buffer assertion.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"`A"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">125
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents an end of buffer assertion.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"'z"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">126
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The newline character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"\n"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">127
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The comma separator.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">","
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">128
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the bell character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"a"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">129
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the form feed character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"f"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">130
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the newline character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"n"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">131
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the carriage return character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"r"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">132
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the tab character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"t"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">133
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the vertical tab character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"v"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">134
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the start of a hexadecimal character constant.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"x"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">135
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the start of an ASCII escape character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"c"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">136
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The colon character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">":"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">137
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The equals character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"="
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">138
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the ASCII escape character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"e"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">139
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents any lower case character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"l"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">140
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents any non-lower case character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"L"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">141
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents any upper case character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"u"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">142
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents any non-upper case character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"U"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">143
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents any space character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"s"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">144
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents any non-space character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"S"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">145
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents any digit character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"d"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">146
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents any non-digit character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"D"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">147
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the end quote operator.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"E"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">148
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the start quote operator.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"Q"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">149
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents a Unicode combining character sequence.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"X"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">150
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents any single character.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"C"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">151
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents end of buffer operator.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"Z"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="21%">152
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character which when preceded by an escape
|
|
character represents the continuation assertion.
|
|
</TD>
|
|
<TD vAlign="top" width="29%">"G"
|
|
</TD>
|
|
<TD vAlign="top" width="9%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD> </TD>
|
|
<TD>153</TD>
|
|
<TD>The character which when preceeded by (? indicates a zero width negated
|
|
forward lookahead assert.</TD>
|
|
<TD>!</TD>
|
|
<TD> </TD>
|
|
</TR>
|
|
</TABLE>
|
|
</P>
|
|
<P><BR>
|
|
|
|
</P>
|
|
<P>Custom error messages are loaded as follows:
|
|
<BR>
|
|
|
|
</P>
|
|
<P>
|
|
<TABLE id="Table3" cellSpacing="0" cellPadding="7" width="624" border="0">
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">Message ID
|
|
</TD>
|
|
<TD vAlign="top" width="32%">Error message ID
|
|
</TD>
|
|
<TD vAlign="top" width="31%">Default string
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">201
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_NOMATCH
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"No match"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">202
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_BADPAT
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Invalid regular expression"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">203
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_ECOLLATE
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Invalid collation character"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">204
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_ECTYPE
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Invalid character class name"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">205
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_EESCAPE
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Trailing backslash"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">206
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_ESUBREG
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Invalid back reference"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">207
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_EBRACK
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Unmatched [ or [^"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">208
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_EPAREN
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Unmatched ( or \\("
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">209
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_EBRACE
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Unmatched \\{"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">210
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_BADBR
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Invalid content of \\{\\}"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">211
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_ERANGE
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Invalid range end"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">212
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_ESPACE
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Memory exhausted"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">213
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_BADRPT
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Invalid preceding regular expression"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">214
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_EEND
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Premature end of regular expression"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">215
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_ESIZE
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Regular expression too big"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">216
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_ERPAREN
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Unmatched ) or \\)"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">217
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_EMPTY
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Empty expression"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">218
|
|
</TD>
|
|
<TD vAlign="top" width="32%">REG_E_UNKNOWN
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"Unknown error"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
</TABLE>
|
|
</P>
|
|
<P><BR>
|
|
|
|
</P>
|
|
<P>Custom character class names are loaded as followed:
|
|
<BR>
|
|
|
|
</P>
|
|
<P>
|
|
<TABLE id="Table4" cellSpacing="0" cellPadding="7" width="624" border="0">
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">Message ID
|
|
</TD>
|
|
<TD vAlign="top" width="32%">Description
|
|
</TD>
|
|
<TD vAlign="top" width="31%">Equivalent default class name
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">300
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character class name for alphanumeric characters.
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"alnum"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">301
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character class name for alphabetic characters.
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"alpha"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">302
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character class name for control characters.
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"cntrl"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">303
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character class name for digit characters.
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"digit"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">304
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character class name for graphics characters.
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"graph"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">305
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character class name for lower case characters.
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"lower"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">306
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character class name for printable characters.
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"print"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">307
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character class name for punctuation characters.
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"punct"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">308
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character class name for space characters.
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"space"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">309
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character class name for upper case characters.
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"upper"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">310
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character class name for hexadecimal characters.
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"xdigit"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">311
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character class name for blank characters.
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"blank"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">312
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character class name for word characters.
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"word"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
<TR>
|
|
<TD vAlign="top" width="8%"> </TD>
|
|
<TD vAlign="top" width="22%">313
|
|
</TD>
|
|
<TD vAlign="top" width="32%">The character class name for Unicode characters.
|
|
</TD>
|
|
<TD vAlign="top" width="31%">"unicode"
|
|
</TD>
|
|
<TD vAlign="top" width="7%"> </TD>
|
|
</TR>
|
|
</TABLE>
|
|
</P>
|
|
<P><BR>
|
|
|
|
</P>
|
|
<P>Finally, custom collating element names are loaded starting from message id
|
|
400, and terminating when the first load thereafter fails. Each message looks
|
|
something like: "tagname string" where <I>tagname</I> is the name used inside
|
|
[[.tagname.]] and <I>string</I> is the actual text of the collating element.
|
|
Note that the value of collating element [[.zero.]] is used for the conversion
|
|
of strings to numbers - if you replace this with another value then that will
|
|
be used for string parsing - for example use the Unicode character 0x0660 for
|
|
[[.zero.]] if you want to use Unicode Arabic-Indic digits in your regular
|
|
expressions in place of Latin digits.
|
|
</P>
|
|
<P>
|
|
Note that the POSIX defined names for character classes and collating elements
|
|
are always available - even if custom names are defined, in contrast, custom
|
|
error messages, and custom syntax messages replace the default ones.
|
|
<P>
|
|
<HR>
|
|
<P></P>
|
|
<p>Revised
|
|
<!--webbot bot="Timestamp" S-Type="EDITED" S-Format="%d %B, %Y" startspan -->
|
|
11 April 2003
|
|
<!--webbot bot="Timestamp" endspan i-checksum="39359" -->
|
|
</p>
|
|
<P><I>� Copyright <a href="mailto:jm@regex.fsnet.co.uk">John Maddock</a> 1998-<!--webbot bot="Timestamp" S-Type="EDITED" S-Format="%Y" startspan --> 2003<!--webbot bot="Timestamp" endspan i-checksum="39359" --></I></P>
|
|
<P align="left"><I>Permission to use, copy, modify, distribute and sell this software
|
|
and its documentation for any purpose is hereby granted without fee, provided
|
|
that the above copyright notice appear in all copies and that both that
|
|
copyright notice and this permission notice appear in supporting documentation.
|
|
Dr John Maddock makes no representations about the suitability of this software
|
|
for any purpose. It is provided "as is" without express or implied warranty.</I></P>
|
|
</body>
|
|
</html>
|