diff options
| author | Ulrich Drepper <drepper@redhat.com> | 1997-08-20 03:53:21 +0000 |
|---|---|---|
| committer | Ulrich Drepper <drepper@redhat.com> | 1997-08-20 03:53:21 +0000 |
| commit | 40a55d2054ef9d9085c4e6b0157c5d30608736fc (patch) | |
| tree | 7ccdbae13abe1988c2ce9af0f936daac639ba1b5 /manual | |
| parent | ca34d7a7c77fa8dc823fd05439870e5a96610c5a (diff) | |
| download | glibc-40a55d2054ef9d9085c4e6b0157c5d30608736fc.tar.xz glibc-40a55d2054ef9d9085c4e6b0157c5d30608736fc.zip | |
Update.cvs/libc-ud-970819
1997-08-20 05:30 Ulrich Drepper <drepper@cygnus.com>
* catgets/catgets.c (catclose): Use __munmap instead of munmap.
* catgets/gencat.c (read_input_file): Fix typo.
* dirent/dirent.h: Make seekdir and telldir available for __USE_XOPEN.
* elf/dl-load.c: Fix case of missing DT_RPATH in object which gets
executed (e.g., when it is a static binary).
* intl/bindtextdomain.c: Use strdup in glibc. Correct comment.
* intl/dcgettext.c: Likewise.
* intl/dgettext.c: Likewise.
* intl/explodename.c: Likewise.
* intl/finddomain.c: Likewise.
* intl/gettext.c: Likewise.
* intl/gettext.h: Likewise.
* intl/hash-string.h: Likewise.
* intl/l10nflist.c: Likewise.
* intl/libintl.h: Likewise.
* intl/loadinfo.h: Likewise.
* intl/loadmsgcat.c: Likewise.
* intl/localealias.c: Likewise.
* intl/textdomain.c: Likewise.
Unify libio sources with code in libg++.
* libio/fcloseall.c: Update and reformat copyright. Protect use
of weak_alias. Use _IO_* thread macros instead of __libc_*.
* libio/feof.c: Likewise.
* libio/feof_u.c: Likewise.
* libio/ferror.c: Likewise.
* libio/ferror_u.c: Likewise.
* libio/fgetc.c: Likewise.
* libio/filedoalloc.c: Likewise.
* libio/fileno.c: Likewise.
* libio/fileops.c: Likewise.
* libio/fputc.c: Likewise.
* libio/fputc_u.c: Likewise.
* libio/freopen.c: Likewise.
* libio/fseek.c: Likewise.
* libio/genops.c: Likewise.
* libio/getc.c: Likewise.
* libio/getc_u.c: Likewise.
* libio/getchar.c: Likewise.
* libio/getchar_u.c: Likewise.
* libio/iofclose.c: Likewise.
* libio/iofdopen.c: Likewise.
* libio/iofflush.c: Likewise.
* libio/iofflush_u.c: Likewise.
* libio/iofgetpos.c: Likewise.
* libio/iofgets.c: Likewise.
* libio/iofopen.c: Likewise.
* libio/iofopncook.c: Likewise.
* libio/iofprintf.c: Likewise.
* libio/iofputs.c: Likewise.
* libio/iofread.c: Likewise.
* libio/iofsetpos.c: Likewise.
* libio/ioftell.c: Likewise.
* libio/iofwrite.c: Likewise.
* libio/iogetdelim.c: Likewise.
* libio/iogetline.c: Likewise.
* libio/iogets.c: Likewise.
* libio/iopadn.c: Likewise.
* libio/iopopen.c: Likewise.
* libio/ioputs.c: Likewise.
* libio/ioseekoff.c: Likewise.
* libio/ioseekpos.c: Likewise.
* libio/iosetbuffer.c: Likewise.
* libio/iosetvbuf.c: Likewise.
* libio/iosprintf.c: Likewise.
* libio/ioungetc.c: Likewise.
* libio/iovdprintf.c: Likewise.
* libio/iovsprintf.c: Likewise.
* libio/iovsscanf.c: Likewise.
* libio/libio.h: Likewise.
* libio/libioP.h: Likewise.
* libio/obprintf.c: Likewise.
* libio/pclose.c: Likewise.
* libio/peekc.c: Likewise.
* libio/putc.c: Likewise.
* libio/putchar.c: Likewise.
* libio/rewind.c: Likewise.
* libio/setbuf.c: Likewise.
* libio/setlinebuf.c: Likewise.
* libio/stdfiles.c: Likewise.
* libio/stdio.c: Likewise.
* libio/strfile.h: Likewise.
* libio/strops.c: Likewise.
* libio/vasprintf.c: Likewise.
* libio/vscanf.c: Likewise.
* libio/vsnprintf.c: Likewise.
* manual/libc.texinfo: Add menu entries for chapter on message
translation.
* manual/locale.texi: Correct next entry in @node for new chapter.
* manual/search.texi: Likewise for previous link.
* manual/message.texi: New file.
* manual/startup.texi: Document LC_ALL, LC_MESSAGES, NLSPATH,
setenv, unsetenv, and clearenv.
* manual/string.texi: Fix typos. Patch by Jim Meyering.
* math/Makefile (test-longdouble-yes): Enable. We want long double
tests now.
Crusade against strcat.
* nis/nss_nisplus/nisplus-publickey.c: Remove uses of strcat.
* stdlib/canonicalize.c: Likewise.
* posix/glob.h: Define __const if necessary. Use __const in all
prototypes.
* sysdeps/generic/stpcpy.c: Use K&R form to allow use in other
GNU packages.
* posix/wordexp.c: Completely reworked buffer handling for much
better performance. Patch by Tim Waugh.
* socket/sys/sochet.h (getpeername): Fix type of LEN parameter,
it must be socklen_t.
* sysdeps/libm-i387/e_remainder.S: Pretty print.
* sysdeps/libm-i387/e_remainderf.S: Likewise.
* sysdeps/libm-i387/e_remainderl.S: Pop extra value for FPU stack.
* sysdeps/libm-i387/s_cexp.S: Little optimization.
* sysdeps/libm-i387/s_cexpl.S: Likewise.
* sysdep/libm-ieee754/s_csinhl.c: Include <fenv.h>.
1997-08-18 15:21 Ulrich Drepper <drepper@cygnus.com>
* sysdeps/unix/sysv/linux/if_index.c (if_nameindex): Fix memory leak
in cleanup code.
1997-08-17 Paul Eggert <eggert@twinsun.com>
* tzset.c (__tzset_internal): Fix memory leak when the user
specifies a TZ value that uses a default rule file.
Do not assume US DST rules when the user specifies
that there is no DST.
1997-08-10 19:17 Philip Blundell <Philip.Blundell@pobox.com>
* inet/getnameinfo.c: Tidy up.
* sysdeps/posix/getaddrinfo.c: Likewise.
* sysdeps/unix/sysv/linux/if_index.c (if_nametoindex): Return 0 if
using stub code.
(if_indextoname): Use SIOGIFNAME ioctl if the kernel supports it.
(if_nameindex): Use alloca() rather than malloc(); use
SIOCGIFCOUNT ioctl if the kernel supports it.
1997-08-16 Andreas Schwab <schwab@issan.informatik.uni-dortmund.de>
* sysdeps/unix/sysv/linux/sys/mount.h: Remove the IS_* macros,
they operate on internal kernel structures and have no place in a
user header.
1997-08-16 Andreas Schwab <schwab@issan.informatik.uni-dortmund.de>
* Makerules (lib%.so): Depend on $(+preinit) and $(+postinit).
(build-shlib): Filter them out of $^.
1997-08-15 Andreas Schwab <schwab@issan.informatik.uni-dortmund.de>
* elf/dl-error.c (_dl_signal_error): Fix error message.
1997-08-16 04:06 Ulrich Drepper <drepper@cygnus.com>
* assert/assert.h [__USE_GNU]: Undefine assert_perror.
Reported by Theodore C. Belding <Ted.Belding@umich.edu>.
1997-08-13 Andreas Schwab <schwab@issan.informatik.uni-dortmund.de>
* Makeconfig: Change object suffixes from *.[spgb]o to *.o[spgb]
to avoid conflict with PO files.
* Makerules: Likewise.
* Rules: Likewise.
* elf/Makefile: Likewise.
* extra-lib.mk: Likewise.
* gmon/Makefile: Likewise.
* nis/Makefile: Likewise.
* nss/Makefile: Likewise.
* resolv/Makefile: Likewise.
* rpm/Makefile: Likewise.
* sunrpc/Makefile: Likewise.
* sysdeps/sparc/elf/Makefile: Likewise.
* sysdeps/sparc64/elf/Makefile: Likewise.
* sysdeps/unix/sysv/linux/sparc/Makefile: Likewise.
(ASFLAGS-.os): Renamed from as-FLAGS.os.
Diffstat (limited to 'manual')
| -rw-r--r-- | manual/libc.texinfo | 8 | ||||
| -rw-r--r-- | manual/locale.texi | 2 | ||||
| -rw-r--r-- | manual/message.texi | 1185 | ||||
| -rw-r--r-- | manual/search.texi | 2 | ||||
| -rw-r--r-- | manual/startup.texi | 85 | ||||
| -rw-r--r-- | manual/string.texi | 4 |
6 files changed, 1276 insertions, 10 deletions
diff --git a/manual/libc.texinfo b/manual/libc.texinfo index cb1769fec7..6a936fdbe4 100644 --- a/manual/libc.texinfo +++ b/manual/libc.texinfo @@ -122,6 +122,8 @@ of the GNU C Library. * Extended Characters:: Support for extended character sets. * Locales:: The country and language can affect the behavior of library functions. +* Message Translation:: How to make the program speak the users + language. * Searching and Sorting:: General searching and sorting functions. * Pattern Matching:: Matching wildcards and regular expressions, and shell-style ``word expansion''. @@ -314,6 +316,11 @@ Locales and Internationalization * Standard Locales:: Locale names available on all systems. * Numeric Formatting:: How to format numbers for the chosen locale. +Message Translation + +* Message catalogs a la X/Open:: The @code{catgets} family of functions. +* The Uniforum approach:: The @code{gettext} family of functions. + Searching and Sorting * Comparison Functions:: Defining how to compare two objects. @@ -975,6 +982,7 @@ Porting the GNU C Library @include time.texi @include mbyte.texi @include locale.texi +@include message.texi @include setjmp.texi @include signal.texi @include startup.texi diff --git a/manual/locale.texi b/manual/locale.texi index 1866c66fb7..dfc9117176 100644 --- a/manual/locale.texi +++ b/manual/locale.texi @@ -1,4 +1,4 @@ -@node Locales, Searching and Sorting, Extended Characters, Top +@node Locales, Message Translation, Extended Characters, Top @chapter Locales and Internationalization Different countries and cultures have varying conventions for how to diff --git a/manual/message.texi b/manual/message.texi new file mode 100644 index 0000000000..7640e21acf --- /dev/null +++ b/manual/message.texi @@ -0,0 +1,1185 @@ +@node Message Translation +@chapter Message Translation + +The program's interface with the human should be designed in a way to +ease the human the task. One of the possibilities is to use messages in +whatever language the user prefers. + +Printing messages in different languages can be implemented in different +ways. One could add all the different languages in the source code and +add among the variants every time a message has to be printed. This is +certainly no good solution since extending the set of languages is +difficult (the code must be changed) and the code itself can become +really big with dozens of message sets. + +A better solution is to keep the message sets for each language are kept +in separate files which are loaded at runtime depending on the language +selection of the user. + +The GNU C Library provides two different sets of functions to support +message translation. The problem is that neither of the interfaces is +officially defined by the POSIX standard. The @code{catgets} family of +functions is defined in the X/Open standard but this is drived from +industry decisions and therefore not necessarily is based on reasinable +decisions. + +As mentioned above the message catalog handling provides easy +extendibility by using external data files which contain the message +translations. I.e., these files contain for each of the messages used +in the program a translation for the appropriate language. So the tasks +of the message handling functions functions are + +@itemize @bullet +@item +locate the external data file with the appropriate translations. +@item +load the data and make it possible to address the messages +@item +map a given key to the translated message +@end itemize + +The two approaches mainly differ in the implementation of this last +step. The design decisions made for this influences the whole rest. + +@menu +* Message catalogs a la X/Open:: The @code{catgets} family of functions. +* The Uniforum approach:: The @code{gettext} family of functions. +@end menu + + +@node Message catalogs a la X/Open +@section X/Open Message Catalog Handling + +The @code{catgets} functions are based on the simple scheme: + +@quotation +Associate every message to translate in the source code with a unique +identifier. To retrieve a message from a catalog file solely the +identifier is used. +@end quotation + +This means for the author of the program that s/he will have to make +sure the meaning of the identifier in the program code and in the +message catalogs are always the same. + +Before a message can be translated the catalog file must be located. +The user of the program must be able to guide the responsible function +to find whatever catalog the user wants. This is separated from what +the programmer had in mind. + +All the types, constants and funtions for the @code{catgets} functions +are defined/declared in the @file{nl_types.h} header file. + +@menu +* The catgets Functions:: The @code{catgets} function family. +* The message catalog files:: Format of the message catalog files. +* The gencat program:: How to generate message catalogs files which + can be used by the functions. +* Common Usage:: How to use the @code{catgets} interface. +@end menu + + +@node The catgets Functions +@subsection The @code{catgets} function family + +@comment nl_types.h +@comment X/Open +@deftypefun nl_catd catopen (const char *@var{cat_name}, int @var{flag}) +The @code{catgets} function tries to locate the message data file names +@var{cat_name} and loads it when found. The return value is of an +opaque type and can be used in calls to the other functions to refer to +this loaded catalog. + +The return value is @code{(nl_catd) -1} in case the function failed and +no catalog was loaded. The global variable @var{errno} contains a code +for the error causing the failure. But even if the function call +succeeded this does not mean that all messages can be translated. + +Locating the catalog file must happen in a way which lets the user of +the program influence the decision. It is up to the user to decide +about the language to use and sometimes it is useful to use alternate +catalog files. All this can be specified by the user by setting some +enviroment variables. + +The first problem is to find out where all the message catalogs are +stored. Every program could have its own place to keep all the +different files but usually the catalog files are grouped by languages +and the catalogs for all programs are kept in the same place. + +@cindex NLSPATH environment variable +To tell the @code{catopen} function where the catalog for the program +can be found the user can set the environment variable @code{NLSPATH} to +a value which describes her/his choice. Since this value must be usable +for different languages and locales it cannot be a simple string. +Instead it is a format string (similar to @code{printf}'s). An example +is + +@smallexample +/usr/share/locale/%L/%N:/usr/share/locale/%L/LC_MESSAGES/%N +@end smallexample + +First one can see that more than one directory can be specified (with +the usual syntax of separating them by colons). The next things to +observe are the format string, @code{%L} and @code{%N} in this case. +The @code{catopen} function knows about several of them and the +replacement for all of them is of course different. + +@table @code +@item %N +This format element is substituted with the name of the catalog file. +This is the value of the @var{cat_name} argument given to +@code{catgets}. + +@item %L +This format element is substituted with the name of the currently +selected locale for translating messages. How this is determined is +explained below. + +@item %l +(This is the lowercase ell.) This format element is substituted with the +language element of the locale name. The string decsribing the selected +locale is expected to have the form +@code{@var{lang}[_@var{terr}[.@var{codeset}]]} and this format uses the +first part @var{lang}. + +@item %t +This format element is substituted by the territory part @var{terr} of +the name of the currently selected locale. See the explanation of the +format above. + +@item %c +This format element is substituted by the codeset part @var{codeset} of +the name of the currently selected locale. See the explanation of the +format above. + +@item %% +Since @code{%} is used in a meta character there must be a way to +express the @code{%} character in the result itself. Using @code{%%} +does this just like it works for @code{printf}. +@end table + + +Using @code{NLSPATH} allows to specify arbitrary directories to be +searched for message catalogs while still allowing different languages +to be used. If the @code{NLSPATH} environment variable is not set the +default value is + +@smallexample +@var{prefix}/share/locale/%L/%N:@var{prefix}/share/locale/%L/LC_MESSAGES/%N +@end smallexample + +@noindent +where @var{prefix} is given to @code{configure} while installing the GNU +C Library (this value is in many cases @code{/usr} or the empty string). + +The remaining problem is to decide which must be used. The value +decides about the substitution of the format elements mentioned above. +First of all the user can specify a path in the message catalog name +(i.e., the name contains a slash character). In this situation the +@code{NLSPATH} environment variable is not used. The catalog must exist +as specified in the program, perhaps relative to the current working +directory. This situation in not desirable and catalogs names never +should be written this way. Beside this, this behaviour is not portable +to all other platforms providing the @code{catgets} interface. + +@cindex LC_ALL environment variable +@cindex LC_MESSAGES environment variable +@cindex LANG environment variable +Otherwise the values of environment variables from the standard +environemtn are examined (@pxref{Standard Environment}). Which +variables are examined is decided by the @var{flag} parameter of +@code{catopen}. If the value is @code{NL_CAT_LOCALE} (which is defined +in @file{nl_types.h}) then the @code{catopen} function examines the +environment variable @code{LC_ALL}, @code{LC_MESSAGES}, and @code{LANG} +in this order. The first variable which is set in the current +environment will be used. + +If @var{flag} is zero only the @code{LANG} environment variable is +examined. This is a left-over from the early days of this function +where the other environment variable were not known. + +In any case the environment variable should have a value of the form +@code{@var{lang}[_@var{terr}[.@var{codeset}]]} as explained above. If +no environment variable is set the @code{"C"} locale is used which +prevents any translation. + +The return value of the function is in any case a valid string. Either +it is a translation from a message catalog or it is the same as the +@var{string} parameter. So a piece of code to decide whether a +translation actually happened must look like this: + +@smallexample +@{ + char *trans = catgets (desc, set, msg, input_string); + if (trans == input_string) + @{ + /* Something went wrong. */ + @} +@} +@end smallexample + +@noindent +When an error occured the global variable @var{errno} is set to + +@table @var +@item EBADF +The catalog does not exist. +@item ENOMSG +The set/message touple does not name an existing element in the +message catalog. +@end table + +While it sometimes can be useful to test for errors programs normally +will avoid any test. If the translation is not available it is no big +problem if the original, untranslated message is printed. Either the +user understands this as well or s/he will look for the reason why the +messages are not translated. +@end deftypefun + +Please note that the currently selected locale does not depend on a call +to the @code{setlocale} function. It is not necessary that the locale +data files for this locale exist and calling @code{setlocale} succeeds. +The @code{catopen} function directly reads the values of the environment +variables. + + +@deftypefun {char *} catgets (nl_catd @var{catalog_desc}, int @var{set}, int @var{message}, const char *@var{string}) +The function @code{catgets} has to be used to access the massage catalog +previously opened using the @code{catopen} function. The +@var{catalog_desc} parameter must be a value previously returned by +@code{catopen}. + +The next two parameters, @var{set} and @var{message}, reflect the +internal organization of the message catalog files. This will be +explained in detail below. For now it is interesting to know that a +catalog can consists of several set and the messages in each thread are +individually numbered using numbers. Neither the set number nor the +message number must be consecutive. They can be arbitrarily chosen. +But each message (unless equal to another one) must have its own unique +pair of set and message number. + +Since it is not guaranteed that the message catalog for the language +selected by the user exists the last parameter @var{string} helps to +handle this case gracefully. If no matching string can be found +@var{string} is returned. This means for the programmer that + +@itemize @bullet +@item +the @var{string} parameters should contain reasonable text (this also +helps to understand the program seems otherwise there would be no hint +on the string which is expected to be returned. +@item +all @var{string} arguments should be written in the same language. +@end itemize +@end deftypefun + +It is somewhat uncomfortable to write a program using the @code{catgets} +functions if no supporting functionality is available. Since each +set/message number touple must be unique the programmer must keep lists +of the messages at the same time the code is written. And the work +between several people working on the same project must be coordinated. +In @ref{Common Usage} we will see some how these problems can be relaxed +a bit. + +@deftypefun int catclose (nl_catd @var{catalog_desc}) +The @code{catclose} function can be used to free the resources +associated with a message catalog which previously was opened by a call +to @code{catopen}. If the resources can be successfully freed the +function returns @code{0}. Otherwise it return @code{@minus{}1} and the +global variable @var{errno} is set. Errors can occur if the catalog +descriptor @var{catalog_desc} is not valid in which case @var{errno} is +set to @code{EBADF}. +@end deftypefun + + +@node The message catalog files +@subsection Format of the message catalog files + +The only reasonable way the translate all the messages of a function and +store the result in a message catalog file which can be read by the +@code{catopen} function is to write all the message text to the +translator and let her/him translate them all. I.e., we must have a +file with entries which associate the set/message touple with a specific +translation. This file format is specified in the X/Open standard and +is as follows: + +@itemize @bullet +@item +Lines containing only whitespace characters or empty lines are ignored. + +@item +Lines which contain as the first non-whitespace character a @code{$} +followed by a whitespace character are comment and are also ignored. + +@item +If a line contains as the first non-whitespace characters the sequence +@code{$set} followed by a whitespace character an additional argument +is required to follow. This argument can either be: + +@itemize @minus +@item +a number. In this case the value of this number determines the set +to which the following messages are added. + +@item +an identifier consisting of alphanumeric characters plus the underscore +character. In this case the set get automatically a number assigned. +This value is one added to the largest set number which so far appeared. + +How to use the symbolic names is explained in section @ref{Common Usage}. + +It is an error if a symbol name appears more than once. All following +messages are placed in a set with this number. +@end itemize + +@item +If a line contains as the first non-whitespace characters the sequence +@code{$delset} followed by a whitespace character an additional argument +is required to follow. This argument can either be: + +@itemize @minus +@item +a number. In this case the value of this number determines the set +which will be deleted. + +@item +an identifier consisting of alphanumeric characters plus the underscore +character. This symbolic identifier must match a name for a set which +previously was defined. It is an error if the name is unknown. +@end itemize + +In both cases all messages in the specified set will be removed. They +will not appear in the output. But if this set is later again selected +with a @code{$set} command again messages could be added and these +messages will appear in the output. + +@item +If a line contains after leading whitespaces the sequence +@code{$quote}, the quoting character used for this input file is +changed to the first non-whitespace character following the +@code{$quote}. If no non-whitespace character is present before the +line ends quoting is disable. + +By default no quoting character is used. In this mode strings are +terminated with the first unescaped line break. If there is a +@code{$quote} sequence present newline need not be escaped. Instead a +string is terminated with the first unescaped appearence of the quote +character. + +A common usage of this feature would be to set the quote character to +@code{"}. Then any appearence of the @code{"} in the strings must +be escaped using the backslash (i.e., @code{\"} must be written). + +@item +Any other line must start with a number or an alphanumeric identifier +(with the underscore character included). The following characters +(starting at the first non-whitespace character) will form the string +which gets associated with the currently selected set and the message +number represented by the number and identifier respectively. + +If the start of the line is a number the message number is obvious. It +is an error if the same message number already appeared for this set. + +If the leading token was an identifier the message number gets +automatically assigned. The value is the current maximum messages +number for this set plus one. It is an error if the identifier was +already used for a message in this set. It is ok to reuse the +identifier for a message in another thread. How to use the symbolic +identifiers will be explained below (@pxref{Common Usage}). There is +one limitation with the identifier: it must not be @code{Set}. The +reason will be explained below. + +Please note that you must use a quoting character if a message contains +leading whitespace. Since one cannot guarantee this never happens it is +probably a good idea to always use quoting. + +The text of the messages can contain escape characters. The usual bunch +of characters known from the @w{ISO C} language are recognized +(@code{\n}, @code{\t}, @code{\v}, @code{\b}, @code{\r}, @code{\f}, +@code{\\}, and @code{\@var{nnn}}, where @var{nnn} is the octal coding of +a character code). +@end itemize + +@strong{Important:} The handling of identifiers instead of numbers for +the set and messages is a GNU extension. Systems strictly following the +X/Open specification do not have this feature. An example for a message +catalog file is this: + +@smallexample +$ This is a leading comment. +$quote " + +$set SetOne +1 Message with ID 1. +two " Message with ID \"two\", which gets the value 2 assigned" + +$set SetTwo +$ Since the last set got the nubmer 1 assigned this set has number 2. +4000 "The numbers can be arbitrary, they need not start at one." +@end smallexample + +This small example shows various aspects: +@itemize @bullet +@item +Lines 1 and 9 are comments since they start with @code{$} followed by +a whitespace. +@item +The quoting character is set to @code{"}. Otherwise the quotes in the |
