7. Lexical Analysis and Preprocessing#
7.1. Source Line Data Structure#
Before describing the routines involved in lexical analysis and preprocessing,
it is necessary to discuss the data structure that represents the current
source line. This data structure is used throughout the lexical routines, and
its form is dictated by the types of processing that must be done in the
lexical and macro routines. The routines and data structures described in this
section are defined in lexical.c and lexical.h.
7.1.1. The Current Source Line#
A logical source line, in the ANSI C standard (2.1.2.2), is what results after
trigraphs have been replaced by the corresponding characters (e.g., ??= by
#) and line splices have been done (i.e., each line ending with a backslash
is spliced with the following line). The routine read_logical_source_line
does this processing, and the global array curr_source_line contains the
result. That is, in curr_source_line trigraphs have already been replaced,
and several physical source lines may have been spliced into one logical source
line. If any such transformations were made to produce the current source
line, the global variable orig_line_modif_list points at a list of entries
that describe all the transformations made, in a form such that
- the original source lines can be reconstructed from
curr_source_line, if that is necessary (e.g., for error-reporting purposes); and - from a location within
curr_source_line, one can determine the file name, line number, and column from which that text came, for error-reporting purposes.
7.1.2. Modifications to the Current Source Line#
When changes to the current logical source line are made because of
preprocessing, they are handled in a different way: curr_source_line is not
changed (much); instead, the global variable source_line_modif_list points
to a list of entries that indicate changes that have been made logically to
curr_source_line. Each of these changes is a deletion of some number of
characters at a given location, and an insertion of some number of characters
(possibly none) at that location. These entries represent changes due to
deletion of comments and replacement of macro invocations by the expansions of
those macros. (Comments need not be deleted, even logically, if preprocessing
output is not being produced. In pcc mode, comments are deleted entirely,
but in other modes they are replaced by one blank.)
The two sets of changes are handled differently, and one might ask why the
trigraphs and line splices are actually rendered in curr_source_line (with
instructions on undoing them), while the macro changes are not rendered (and
the instructions are for doing them rather than undoing them). The reason is
that the first set of changes is applicable at the character or intra-token
level, and the second set only applies at the inter-token level. It is a basic
goal of the lexical routines that the current position must be representable as
a single pointer, and that one be able to move forward efficiently in the input
by just incrementing that pointer. Trigraphs and line splices can occur within
tokens, and therefore if they were not rendered in curr_source_line one
would have to check for them in every case where one wants to move forward in
the input. The other modifications, on the other hand, can only occur between
tokens, and therefore they need not be expanded in the source line. They occur
between tokens, and special marker characters in the input stream are used to
indicate their presence. The marker characters cause the termination of
scanning of the previous token, followed (on the next call of get_token) by
interpretation of the marker character to decide where the input goes next.
This makes macro expansion efficient (the source line need not actually be
subjected to insertions and deletions) while preserving the ability to advance
through the input using just one pointer. It also preserves information that
allows one to construct either the original source line or the macro-expanded
line, and to know the origin of any particular part of the resultant line.
The modification entries can apply to the inserted text of other modifications,
so one can have modifications of modifications. Wherever a modification is
made, the original character (in curr_source_line or inserted text) is
replaced by a marker character (ATTENTION_MARKER) so that if that spot is
reached by sequential scanning from preceding text, one is cued to look at the
modification list to find the text that follows.
In extreme cases the data structure can get quite complex, with modifications on modifications. In cases involving real-world macros, however, it remains fairly simple.
ATTENTION_MARKER is a one-character lexical escape sequence. It has to be,
because a macro invocation can be as short as a single character, and the
ATTENTION_MARKER overwrites the beginning of the macro invocation. There
are also two-character lexical escape sequences, which begin with the
LE_ESCAPE character (which is zero). The second character of the escape
sequence is one of the following:
LE_END_OF_LINE, which marks the end of the primary source line.LE_NEWLINE, which represents the newline character at the end of the primary source line.LE_END_OF_INSERTION, which marks the end of the text inserted by a source line modification.LE_END_OF_TOKEN, which separates tokens in macro text (see below).LE_INERT_MACRO, which indicates a macro name that should not be expanded (see below).LE_NULL, which indicates a null (zero) character in the source line.
To ensure that the ATTENTION_MARKER and LE_ESCAPE characters are always
lexical escapes, rather than characters accidentally introduced into the
lexical data structure because of special characters in the input file, the
front end does two rewrites on input:
- The newline character at the ends of input lines is transformed into a lexical escape sequence. This frees up the newline character to be used as
ATTENTION_MARKER. - A null (zero) character in an input line is flagged with an error and replaced by a space. In modes that allow null characters in the source (e.g., gcc mode), the character is replaced by an
LE_NULLescape sequence (and an original line modification is created, so that column numbers can be determined correctly).
The characters newline and null were selected for use as these reserved characters because they are already special in the C language and in character sets in general. No other pair of characters can be guaranteed not to occur in source files, especially when international character encodings such as SJIS are considered.
Note that because LE_ESCAPE is a zero character, the source line and
inserted text cannot be considered to be simple null-terminated strings. Also
note that the ATTENTION_MARKER and LE_ESCAPE characters will naturally
terminate the accumulation of the characters of a token.
The modification entries of both kinds are dynamically allocated, used, and then freed when the next logical source line is about to be read. For each kind, there is an available list (freed entries that are available for reuse), an allocation routine, and a free routine. See
add_orig_line_modif,free_orig_line_modif,add_source_line_modif,free_source_line_modif, andrem_source_line_modif.
The majority (75%?) of source lines have no modifications of any kind (not counting comments), so the code in the lexical routines is optimized to deal efficiently with that case.
Routines assoc_source_line_modif and nested_source_line_modif, along
with the associated macros go_into_insertion and leave_insertion,
handle the low-level tasks of determining where to go next in the input stream
based on the source modifications.
The current position in the input stream in given by curr_char_loc.
Ordinarily, it points to somewhere within curr_source_line, but when macros
are expanded it can point to somewhere in macro_buffer or to a
dynamically-allocated argument.
In the text of macro arguments and the replacement text for macros, the
individual characters have already been tokenized once. To ensure that the
tokenization is done the same way if and when it is done again, the text of
each token in source line modifications is followed by an LE_END_OF_TOKEN
lexical escape sequence. That escape sequence terminates the accumulation of
characters to make up a token, and is then ignored as part of the white-space
skip at the start of the next call of get_token. End-of-token sequences
are usually removed by gen_pp_output_for_curr_line, but an individual
end-of-token will be output as a blank if there is the possibility that two
adjacent tokens might be misparsed without one. See the discussion in
macro.c for further information.
In some obscure cases involving thwarted macro expansion, it is necessary to do
a true insertion into the source line, rather than the usual replacement of
some characters by some others. Fortunately, this only needs to be done at the
beginning of the line, so a simple mechanism suffices: a source line
modification with a NULL location is taken to be an insertion before the first
character of the line. Only one such insertion is allowed per line, and if one
exists, line_start_source_line_modif points to it.
curr_source_line and macro_buffer are actually dynamically-allocated
arrays, and can be expanded (by reallocation) as necessary. Therefore, there
is no limit to the length of source lines or of macros except that imposed by
available memory.
7.2. Lexical Analysis#
“Lexical analysis” includes the reading of source lines and the breaking of
them into lexical tokens. It is also closely tied to preprocessor directives
and to macro processing, since they are invoked at the lexical level, and are
defined in terms of tokens. The source code for lexical processing is in
lexical.c, and associated declarations are in lexical.h.
7.2.1. Getting Tokens#
The primary interface to the lexical routines is the routine get_token,
which fetches and returns the next token of input. When compiling C, these
tokens are the input for syntax analysis. When preprocessing only, the tokens
are fetched (by a loop in cpp_driver in preproc.c), and then discarded;
calling get_token repeatedly forces the reading of input lines and the
necessary preprocessing on each line, and each modified line is then written
out just before the next line is read. do_preprocessing_only and
generate_pp_output will be TRUE in that case. get_token is used in
both cases; in fact, it is used for all fetching of tokens in all contexts.
When get_token is called, it scans the next token of input and returns its
kind. The kind of token is also saved in curr_token. See the enumeration
a_token_kind, in lexical.h, for the list of token kinds (e.g.,
tok_identifier, tok_lparen).
Much of the rest of the processing is dependent on several global switches, as
described in the following paragraphs. Most of these switches are defined in
preproc.h.
If fetch_pp_tokens is TRUE when get_token is called, a preprocessing
token (pp-token) is fetched instead of a regular token. Constants are not
converted (and numeric ones are returned as pp-numbers); adjacent string
literals are not concatenated; keywords are not recognized; and errors are not
issued for malformed tokens. Also, start_of_curr_token and
end_of_curr_token are set to point to the beginning and end of the current
token, and len_of_curr_token is set to its length.
If expand_macros is TRUE, preprocessing macros are expanded as they are
scanned. The caller sees only the tokens after expansion. If a macro name is
preceded by an LE_INERT_MACRO escape sequence, the macro is not expanded
and the identifier is returned to the caller with curr_token_is_inert_macro
set to TRUE.
When fetch_pp_tokens is FALSE, do_string_literal_concatenation controls
whether adjacent string literals are concatenated.
If in_preprocessing_directive is TRUE, the definition of white space is
changed to that for within preprocessing directives; i.e., keywords are not
recognized, and newline, #, and ## are recognized and returned as
tokens. However, if caching_pragma_tokens is TRUE, a pragma is being
recoded as a token cache, and this requires disabling some of the processing
that is normally done when in_preprocessing_directive is TRUE –
identifiers will be looked up and # and ## are not returned as tokens.
And if recognize_keywords_in_pragma is also set, keywords are looked up and
returned as keyword tokens rather than as identifier tokens. (This is required
for dealing with #pragma directives involving C or C++ constructs that must
be scanned by the compiler’s normal scanning routines; for instance, it is used
for processing the pragmas that control template instantiation.)
If in_pp_if_expression is TRUE (indicating that we are inside a
preprocessing #if expression), integer constants will get an implicit L
suffix, and undefined identifiers will be returned as the integer constant
0L.
If exp_header_name is TRUE, then "..." will be scanned as a header name
for an #include (tok_header_name). Other tokens will be processed
normally. exp_system_header_name is similar, for header names of the form
<...>.
If exp_digit_sequence is TRUE (indicating a string of decimal digits) then
the digits will be scanned as a tok_digit_sequence (used in the #line
directive). Other tokens will be processed normally.
For tokens that are literal constants, the struct const_for_curr_token is
set to the value of the constant (if fetch_pp_tokens is FALSE). This
global variable is also set for the literal portion of user-defined literals
(when user_defined_literals_enabled is TRUE), along with
ud_lit_op_sym_for_cur_token, which designates the symbol for the literal
operator or literal operator template chosen to produce the value of the
user-defined literal. When ud_lit_op_sym_for_curr_token designates a raw
literal operator or a literal operator template, const_for_curr_token is a
ck_string constant whose value is the actual spelling of the literal part of
the toke; otherwise, it is the value of that literal part.
For tokens that are identifiers, the corresponding symbol header is looked up
in the symbol table and locator_for_curr_id is set accordingly. For
user-defined literal tokens, locator_for_curr_id is set to designate the
literal-operator-id of the literal operator or literal operator template
denoted by the literal’s ud-suffix. It can be used later to refer to the
current definition of the identifier or operator (if there is one) or to enter
a new definition for the identifier or operator.
The routine get_token is basically one large switch statement that fans
out based on the current character of input. The various cases are:
- For operators and similar special-character tokens,
get_tokenlooks at the next character or two and selects the appropriate token. - For the lexical escape sequences (the ATTENTION_MARKER character or the two-character LE_ESCAPE sequences),
skip_white_spaceis called. These characters are not all actually white space, but it simplifies processing to let that routine handle them. All reads of new source lines are requested byskip_white_space. - For numbers,
scan_numberis called (it handles decimal, octal, and hexadecimal integer literals, as well as decimal and hexadecimal fixed-point and float literals).scan_numberwill determine the kind of number, accumulate its characters, and, if appropriate, call the proper conversion routine inliterals.cto convert the literal constant string to internal form inconst_for_curr_token. It also handles recognition of numeric user-defined literals and, when one is recognized, setting ud_lit_op_sym_for_curr_token. - For identifiers, the characters are accumulated and
find_symbolis called. This yields a list of symbols with the given name. The list is then searched for a definition that is a macro or keyword. If the identifier is a macro and if macro expansion is required,macro_invocation(inmacro.c) is called. Note that because this look-up finds both macros and other uses of identifiers, a symbol only has to be looked up once; it does not have to be looked up once to see whether or not it is a macro, and then again later to see if it is something else. - For quoted tokens (character constants and string literals, and the wide versions thereof), an appropriate scanning routine is called, which in turn calls
accum_quoted_string. The latter accumulates the characters of the string, checking for escapes and the like. When a string literal or wide string literal is scanned outside of preprocessing mode,get_tokenwill callconcat_adjacent_string_literals, which will check if the next token is another string literal. If so, the adjacent strings will be concatenated into a single string literal token. scan_char_constant recognizes user-defined character literals and sets ud_lit_op_sym_for_curr_token, as appropriate. In order to handle merging of suffixes across concatenated string literals, the similar task for user-defined string literals is handled by concat_adjacent_string_literals, except when fetch_pp_tokens is TRUE; in that case, get_token performs the requisite processing, as string literals are not concatenated in that mode. - When a
#appears as the first non-white-space character on a line, a call topp_directive(inpreproc.c) is made to process a preprocessing directive.
One very important requirement is that get_token be fast. About 30 percent
of total time time in the front end is likely to be spent in get_token and
its subordinate routines. Every effort has been made to make those routines
fast, and in particular, to avoid unnecessary subroutine calls. As an example,
simple white space (blanks and horizontal tabs) is skipped in get_token
itself, rather than incurring the overhead of calling skip_white_space.
The handling for pp-tokens and normal tokens differs in a number of ways, but
the most fundamental of those is the information returned to the caller. The
kind of token is returned in both cases. However, for identifiers and literal
constants, where additional information must be returned to precisely specify
the token, the returned value in the pp-token case is the actual characters of
the token, and in the normal token case it is something one step removed from
that: a symbol locator for identifiers, or a constant entry for constants.
This is useful, but it’s also necessary – when scanning normal tokens, and in
particular identifiers and literal constants, there are some cases where it is
necessary to skip forward to the next token to see if it has some effect on the
current one. For example, after scanning a string literal, one must look ahead
to see if the next token is another string literal, and if so, one concatenates
the two. Since the next token may be on another line, it’s clear that the
original string for the token in the current source line may be gone by the
time the next token is reached. Therefore, the locator or constant entry form
is necessary even inside of get_token.
7.2.2. Token Lookahead and Caching#
Sometimes syntactic analysis requires lookahead – it requires the fetching of one or more tokens past the current position to discriminate between two possible parses.
This comes up even in C, although the lookahead there is limited to a single
token. For example, when a statement begins with an identifier, one must look
past the identifier to see whether or not the next token is a “:” – if so,
the identifier is a label. (The routine next_token provides this
capability.)
In C++, lookahead is needed in many more contexts, and is not limited to one-token lookahead. In addition, there are some language constructs that must be saved in some lexical form when they appear, and then parsed later on the basis of the saved information. The most obvious example is the bodies of member functions declared inside a class: their name binding cannot be done until the entire class declaration has been scanned, so the easiest implementation is to save the member function bodies and then parse them when the closing brace of the class is reached.
A construct called a token cache supports all these needs for lookahead, as well as recording the tokens of template definitions and other syntactic elements for delayed use. a_token_cache is a data structure that represents a sequence of tokens. The tokens in a cache are represented by a list of shared tokens (a_shared_token), each one describing an immutable token state (this includes not just the token kind, but also any additional information required to fully specify the token, e.g., a symbol locator for an identifier or a constant entry for a literal constant). A similar data structure, a_reusable_token_cache, is used for the tokens of a template definition or delayed construct.
To build a nonreusable token cache, one declares a variable of type a_token_cache. Then one reads through tokens in the usual way, calling get_token to fetch each token, and then calls cache_curr_token before advancing to the next. When enough lookahead or caching has been done, one can put all the tokens in the token cache back on the front of the input stream by calling rescan_cached_tokens.
get_token does not read directly from nonreusable token caches (although it does do so with reusable token caches; see below). Instead, rescan_cached_tokens adds the current token to the end of the cache and then pushes the cached tokens in LIFO order onto the global cached_token_rescan_stack (implemented as a stack instead of a list for efficiency reasons), and finally calls get_token to fetch what was the first token of the cache from the cached token rescan stack as the current token. As long as the global rescan stack is non-empty, get_token will fetch a token from the rescan stack (by calling get_token_from_cached_token_rescan_stack) instead of scanning a new token from the current source line.
next_token is a simple user of the token caching facility; it caches the
current token, fetches the next and remembers its kind, then pushes the cache
back on the input stream to restore the original token (and also the next one,
behind it).
unget_token puts the current token back on the front of the input stream,
so that in effect afterwards there is no current token (although it doesn’t
actually go so far as to change curr_token). It is used in some unusual
error situations that want to fabricate and insert an identifier token where
one was missing in the source.
Two routines are available to cache sequences of tokens delimited by given tokens:
cache_token_stream_until_matching_tokencache_token_stream_with_coalesce_flag
cache_token_stream_with_coalesce_flag is not intended to be called
directly. Instead, it is called by one of two interface routines:
cache_token_streamcache_token_stream_coalesce_identifiers
cache_token_stream is used to cache tokens until one of a specified set of
stop tokens is found. This is used, for example, when caching the body of a
class or function template when all tokens up to a closing brace should be
cached. In such situations, the only special processing that is needed is to
keep track of paired delimiter tokens (e.g., {}, [], ()) so that,
for example, the closing brace of an inner block is not incorrectly taken as
the end of the template body.
When the stop token could appear within a template argument list, additional
processing must be done to determine whether a “<” that is encountered is
the start of a template argument list or simply a less than sign. In order to
determine this, the identifiers found while the tokens are cached must be
coalesced (the identifier coalescing process is described later in this
chapter). This special kind of token stream caching is done by calling
cache_token_stream_coalesce_identifiers. The process of coalescing an
identifier involves fetching some number of tokens. Because these tokens are
not fetched directly by cache_token_stream_coalesce_identifiers, a
mechanism is needed to get access to those tokens.
begin_caching_of_fetched_tokens is called to request that a cache of
fetched tokens be automatically constructed. Each new token returned by
get_token is added to a special cache. The token stream is then cached by
recording the starting token sequence number, getting each token, and
coalescing each identifier that is encountered until the stop token is found.
Once the starting and ending token sequence numbers are known, the tokens in
that range are extracted from the special cache.
end_caching_of_fetched_tokens is then used to indicate that the automatic
caching of tokens is no longer needed.
As noted above, token caches are also used to represent template definitions and other syntactic constructs that require delayed processing. The tokens that comprise the syntactic element are cached when it is scanned and are rescanned later when needed. While other token caches are rescanned once and discarded, reusable token caches may be rescanned over and over again. Reusable token caches should be constructed by calling the factory function shared_obj<a_token_cache> with reusable set to TRUE (this acts as a sanity check to ensure the token cache is used in the intended way, but they are otherwise built in exactly the same manner as other token caches). The reusable token cache object (of type a_reusable_token_cache) is then constructed from the result.
Various scenarios can cause multiple reusable token caches to be pending rescan at a given time; calling rescan_reusable_cache pushes a reusable token cache onto the global reusable_cache_stack. get_token first rescans tokens from the (nonreusable) cached token rescan stack, as described above. When this stack is exhausted it rescans tokens from the reusable cache at the top of the reusable cache stack by calling get_token_from_reusable_cache_stack. When all of the entries in a reusable cache have been scanned the cached token rescan stack is restored and the entry is popped off of the reusable cache stack. Token scanning then resumes with the next token on the cached token rescan stack.
Reusable and nonreusable token caches may be freely intermixed; template instantiations may occur while scanning tokens from a cache and a template instantiation will usually create new caches, as needed, when doing look-aheads, as described earlier.
7.2.3. Source Positions#
Each time a token is fetched, the global variable pos_curr_token is set to
the source position of the first character of the token. A source position
means a sequence number (which can be converted to a file name and line number)
and a column number from the original source line. The position is determined
by the macro macro_line_loc_to_source_pos. In the simplest and most common
case, a pointer to the starting character of the current token can be converted
easily to a source position by using the current sequence number and
determining the column number by taking the difference between the pointer and
the start of curr_source_line and applying an offset to convert the byte
offset into a logical column number using the macro logical_column_offset
(which can in some cases result in a call of f_logical_column_offset). In
more complicated cases, conv_line_loc_to_source_pos is called, and it uses
the information in both orig_line_modif_list and source_line_modif_list
to determine the source position. The source position of something in a
preprocessing modification is taken to be the position of the first character
of the original source line that contains the macro invocation that leads to
the given modification.
7.2.4. The Input Stack#
open_file_and_push_input_stack is called to enter both the primary source
file and #include files – by fe_init (in fe_init.c) and by
proc_include (in preproc.c), respectively – and it in turn calls
open_file_for_input and push_input_stack to do the work; when end of
file is reached, pop_input_stack is called.
open_file_for_input loops to try each possible include file directory when
attempting to open an include file; it also contains code for the (optional)
implicit inclusion feature of automatic template instantiation, permitting a
series of suffixes to be tried in each directory.
Together, the push/pop routines manage the input_stack, which contains
information about the current nesting of input files. To limit the number of
open files, they will reuse files once a few files have been opened (see
MAX_INCLUDE_FILES_OPEN_AT_ONCE). The storage for input_stack is
dynamically allocated and is reallocated larger as needed.
If the options requesting makefile information or a list of header files
were specified on the command line, do_preprocessing_only will be TRUE but
generate_pp_output will be FALSE. Either list_makefile_dependencies or
list_included_files will be TRUE. push_input_stack writes out the
appropriate information to the preprocessing output file.
The front end recognizes certain idioms used to allow an include file to be
included more than once with all of the includes after the first having no
effect. When one of these idioms is recognized the front end will ignore
subsequent attempts to include the file again. The front end recognizes two
commonly used mechanisms used to allow a file to be included multiple time:
placing the entire contents of a file inside a #ifndef or #ifdef, and
use of the #pragma once directive. A simple state machine is used to
determine whether subsequent includes of a given file may be suppressed. The
state information is maintained in the input stack entry.
suppress_subsequent_include_of_file is called for each file to be included.
If it returns TRUE then the call of push_input_stack is skipped and the
include is suppressed. See the comments that precede find_include_history
for additional information.
7.2.5. Skipping White Space#
skip_white_space is called to skip over any white space and get to the
start of the next token (some simple cases of white space are skipped in
get_token, for efficiency). Besides the usual white-space characters, it
also handles the skipping of comments, calling read_logical_source_line to
fetch the next line of input, and going into and leaving source modifications
upon encountering lexical escape sequences (e.g., ATTENTION_MARKER to go
into a modification, LE_ESCAPE/LE_END_OF_INSERTION to leave a
modification).
skip_white_space tracks the kind of white space it has skipped over
(comment/other) and sets global variable kind_of_white_space_skipped. This
information is useful in the macro processing routines, where white space can
be significant. Note that kind_of_white_space_skipped is not necessarily
set correctly if one calls get_token, since get_token does some
white-space skipping on its own.
When skip_white_space recognizes one of the lint comments
/*ARGSUSED*/,/*NOTREACHED*/,/*VARARGS*/, or/*VARARGSn*/
add_curr_token_pseudo_pragma is called. It adds a pending pragma entry to
the current token pragma list. Later, the pending pragma entry is removed via
a call to extract_specific_pragmas, and processed (by setting global state
information or flags in IL entries). For ARGSUSED and VARARGS, that’s done in
record_lint_argsused_and_varargs_state. For NOTREACHED, it’s done in
check_lint_notreached_state.
7.2.6. Hanging Delete#
If global variable delete_source_from_loc is non-NULL, the text that
extends from that location in the current line to the position immediately
preceding curr_char_loc is to be deleted. This is referred to as a
“hanging delete”, and is used in deleting macro invocations. This flag is
necessary because the complete text of a macro invocation must be deleted (to
be replaced by the expansion of the macro), but the macro invocation may span
several lines, and the macro routines are not in control when ends of lines are
passed in skip_white_space. By setting this flag, the macro routines can
request that the text of a macro invocation be correctly deleted even if it
spills over onto additional lines. When the deletion is handled at line end,
delete_source_from_loc is updated to point to the start of the new source
line so that the hanging delete remains in effect.
7.2.7. Whitespace keywords#
C++/CLI has “whitespace keywords,” which are keywords made up of two words
separated by whitespace but considered a single keyword, for example “refclass”. These are scanned by scan_whitespace_keyword. One interesting
detail is that the characters that make up the token (which can be spread over
multiple lines because of the embedded whitespace) are replaced by a canonical
spelling of the keyword containing a single embedded space, the same way that a
macro is replaced in the input line by its expansion. That allows processing
that deals with tokens on the pp-token level to handle those whitespace keyword
tokens like any others, as a contiguous sequence of characters.
7.2.8. Textual Preprocessing Output#
When compiling C to intermediate language, there is no need to assemble the
actual logical source line that results after preprocessing. All that is
required is that the tokens of that line can be extracted. But when
preprocessing only is being done, the preprocessed line is output from
curr_source_line and the information in source_line_modif_list. [1]
gen_pp_output_for_curr_line assembles and puts out the effective line,
calling gen_pp_line_info to generate line-identification directives when
the file/line of the current line does not immediately follow that of the line
previously output. gen_pp_output_for_curr_line is called by
read_logical_source_line right before a new line is read; it is also called
when the input stack is pushed and popped for the starts and ends of input
files.
There are some special problems in producing the preprocessing output line when
preprocessing directives are passed to output. This must be done for certain
directives (like #pragma) that the preprocessor cannot process. If a
directive of that type contains a multi-line comment, the parts of the
directive preceding and following the comment must be written out on the same
line. In this case, the part preceding the comment will be written without a
terminating newline, and the part following the comment will be written later
on the end of that line.
7.2.9. Reading Logical Source Lines#
Aside from its primary function of reading the next logical source line and
processing trigraphs and line splices, read_logical_source_line also calls
gen_pp_output_for_curr_line to output the previous line (if needed), clears
the modification lists, and handles ends of files. It can be told to stay at
the end of a file (rather than popping out to the enclosing file) on end of
file, if that is desirable (as it is when scanning a comment that is unclosed
at end of file). On the ultimate end of file, it returns a special end-of-file
line (containing just an LE_END_OF_LINE lexical escape sequence) in
curr_source_line, and sets after_end_of_all_source.
If curr_source_line is too small, expand_curr_source_line is called to
reallocate it, and adjust_curr_source_line_structure_after_realloc (in
macro.c) is then called to walk through the data structure associated with
curr_source_line and change any pointers to curr_source_line to point
to the proper addresses within the new allocation for curr_source_line.
7.2.10. Multibyte characters in source code#
When MULTIBYTE_CHARS_IN_SOURCE_SUPPORTED is TRUE, multibyte character
sequences are accepted in comments, string literals, character constants,
identifiers, and header names. Multibyte character sequences are not wide
characters (e.g., Unicode); they are sequences of one or more characters,
appearing in a string of characters of the usual char size, and
representing single logical characters. For example, in the Japanese SJIS
encoding, the sequence 82A9 (in hexadecimal) represents a single Kanji
character: the first character (82) is in a range that indicates it is the
first character of a two-character sequence. Typically, the characters in the
ASCII code set will be accepted as themselves in multibyte character encodings,
i.e., they will be sequences of length one. Since each character sequence in a
multibyte character string need not have the same length (and some encodings
have explicit shift-state-change characters), a multibyte character string can
be understood only by scanning it from beginning to end, advancing over one
logical character at a time.
Moreover, since characters after the first in a multibyte sequence can look
like simple ASCII characters, it is not generally okay to look for an ASCII
character by doing a simple strchr-type search. The lexical routines count
on being able to do this for only two characters: the zero/null character (this
is guaranteed by the C standard; characters after the first in a multibyte
character sequence are not allowed to be zero) and the newline character (this
is effectively guaranteed by the fact that one must be able to read files
containing multibyte character sequences and discern the ends of lines without
tracking the multibyte character state as each character is read). [2]
To advance over logical characters, one needs a routine that will determine the
length of the multibyte character sequence at a given location. This is done
by the macro mbc_length and routine f_mbc_length. In the usual
configuration, that routine is implemented by calling the standard C library
routine mblen. mbc_length and related routines are used to do special
processing in the following places:
- In
read_logical_source_line, the line splice (”\”) and trigraph (”?”) characters are ignored if they appear in the middle of a multibyte character sequence. - In
skip_white_space, the closing character for a C-style comment (”*”) is ignored if it appears in the middle of a multibyte character sequence. (”//” end-of-line comments require no special handling, because the end-of-line marker is rendered as a lexical escape sequence,which cannot appear within a multibyte character sequence.) - In
accum_quoted_string,scan_char_literal,scan_string_literal,conv_single_char, andconv_single_wide_char, the characters inside quoted strings are scanned as multibyte characters as necessary. The routinembc_to_wide_charis used to convert a multibyte character sequence in a wide string literal or wide character constant to the appropriatewchar_tcharacter. In the usual configuration,mbc_to_wide_charis implemented by calling the C library routinembtowc. - In
get_token, the scanning for identifiers usesis_identifier_charto recognize identifier characters beyond the basic character set.make_canonical_identifierhandles encoding the multibyte characters appropriately in the name string for the identifier.
Note that not all source code is scanned as multibyte characters, because that would take too much time. The special scanning is done only in the contexts listed above. For the line-splice and trigraph cases, when the line-read routine sees those characters it goes to the beginning of the line and steps through the characters of the line to find out where the character falls in the multibyte character sequences, but that processing is done only if a potential line splice or trigraph appears. Note also that there are configuration options to turn off the special processing for several of the special characters, for cases where the encoding guarantees no confusion on the particular character.
If the encoding has locking state-change characters (e.g., JIS), each C-style comment and string is assumed to begin in the initial shift state. This follows a requirement stated in 5.2.1.2 of the ANSI/ISO C standard
If UNICODE_SOURCE_SUPPORTED is TRUE, the multibyte encoding selected is the
UTF-8 encoding of Unicode. UTF-16 is also supported, in the little-endian or
big-endian format, and is converted immediately to UTF-8 on input. The type of
source code in a file can be specified by a byte order mark (BOM) at the start
of the file, or by way of the --unicode_source_kind command-line option.
There is also a macro DEFAULT_UNICODE_SOURCE_KIND that can be used to
establish the default for files without a byte order mark.
When UNICODE_SOURCE_SUPPORTED is set to TRUE, identifiers and file name
strings in the front end are represented in UTF-8. For file names, there is
the expectation that the standard system fopen will be able to deal with UTF-8
file names. If that isn’t the case, some host-specific code may have to be
added to fopen_interface. (For Windows-hosted configurations, specific
code is provided which calls the _wfopen function.)
If UTF-8 identifier name strings are somehow undesirable (e.g., a back end or
linker can’t handle them), the macro
IDENTIFIER_STRINGS_ALLOW_MULTIBYTE_CHARS can be set to FALSE, in which case
characters in identifiers outside of the basic character set will be rendered
in the \uxxxx or \Uxxxxxxxx form also used for UCNs. Note,
however, that the encoded form will come out in diagnostic messages that name
an entity in the source.
If DEFAULT_UNICODE_SOURCE_KIND is usk_none, source files without a byte
order mark are assumed to be encoded in ISO-8859-1 (otherwise known as
Latin-1), which matches the first 256 code points in Unicode.
(LOCALE_TO_SET_WHEN_MULTIBYTE_CHARS_ENABLED should be set to a locale that
matches ISO-8859-1.) Such a mode (which is the default for Windows-hosted
versions) is a hybrid mode in which some files are Unicode and some are not,
and in which therefore a character like a u-umlaut would be encoded differently
depending on the kind of source file it appears in. Such characters in
non-Unicode files are converted to UTF-8 (or the identifier encoding \uxxxx) when they appear in identifiers and file names. In hybrid
configurations, characters in textual preprocessing output are always converted
to UTF-8 if DEFAULT_UNICODE_SOURCE_KIND is not usk_none, and to
ISO-8859-1 otherwise. Characters that cannot be represented in ISO-8859-1 in
the latter case are converted to “?”.
When UNICODE_SOURCE_SUPPORTED is TRUE,
NATIVE_MULTIBYTE_CHARS_SUPPORTED_WITH_UNICODE may be set to TRUE to
indicate that non-Unicode files are encoded in some character set other than
ISO-8859-1.This feature requires the ability to translate other character sets
to and from Unicode. Such a facility is provided in newer versions of the
Windows API (Microsoft version 1400 and above) so this feature is enabled when
the front end is built with EDG_WIN32 set to TRUE. In identifiers and file
names, native characters are translated to their UTF-8 equivalent. In wide
strings, native characters are translated to their Unicode equivalent. In
narrow strings, they are left in their original form.
When Microsoft extensions are enabled and EDG_WIN32 is TRUE, this feature
enables support for the Microsoft setlocale pragma, which specifies the
character encoding to be used for the translation to and from Unicode.
7.2.11. Raw Listing Output#
gen_raw_listing_output_for_curr_line is called to write source lines to
f_raw_listing when raw listing information is requested; it calls
gen_expanded_raw_listing_output_for_curr_line to write out the
macro-expanded form of the source line if the source line has some
modifications.
gen_rlisting_line_info writes information about transitions into and out of
include files; it is called by push_input_stack at beginnings of files and
by pop_input_stack at ends of files.
7.2.12. Inserting Text into the Token Stream#
insert_string_into_token_stream may be called to cause a string to be
scanned as tokens that are processed as part of the input stream either before
or after processing the current token. When scanning tokens from the string,
macros are optionally expanded and adjacent string literals are concatenated.
The string cannot contain preprocessing directives, and trigraphs are not
recognized.
This mechanism is used is used to internalize C++ code generated by the C++
metadata reader (see Declarations from C++/CLI Assembly Metadata) in the functions
scan_top_level_metadata_declarations and get_definition_of_class.
7.2.13. Utility Routines#
There are several routines in lexical.c that are utility routines for
syntax analysis, and are not directly related to get_token:
flush_tokensis called on the occurrence of a syntax error, and throws away tokens (by callingget_token) until a token is found that seems an appropriate place to restart. The current stop token set pointed to bycurr_stop_token_stack_entrycontains the set of tokens on which to stop the flush (it is maintained by the macrosadd_stop_token,remove_stop_token, and the routinespush_stop_token_stackandpop_stop_token_stack). Recursive-descent routines place the tokens they expect in the set of stop tokens; then, if an error occurs, the flush will stop at a point that one of the recursive descent routines will recognize.flush_tokensshould not in general be called directly;syntax_errorshould be called to report a syntax error, and it will callflush_tokens.flush_until_matching_tokenis called byflush_tokens(and can be called from elsewhere) to flush from an opening token such as a left parenthesis to the matching closing token.loop_tokenis a utility routine that can be called at the bottom of a loop for a repeated syntactic construct (e.g., a comma-separated list). It tests for a particular token and takes it if it is found.required_tokenis a utility routine that should be called when a particular token is required to appear next. It checks for that token, and issues an error if the token is not found.next_tokenreturns the next token after the current one, without disturbing the current token. This is used in recursive descent parsing routines to look ahead at the next token to decide between two paths through the syntax.next_two_tokensis similar tonext_tokenexcept that if the next token matches a value supplied by the caller the token after the next token is fetched and returned to the caller. This is used when scanning vacuous destructor references in C++ where it is necessary to look for a tilde following a::to recognize a nonclass vacuous destructor reference [3].push_lexical_state_stackandpop_lexical_state_stackare used to enter and exit a new lexical context. These are used, for example, when instantiating a template and when switching from fetching normal tokens to scanning a preprocessing directive.
7.3. Scanning Names in C++#
Unlike C, in which all names are simple identifiers, in C++ the names of
entities are often more complicated. Names may be preceded by one or more
qualifiers such as namespace and class qualifiers (e.g., A::B::i) and
global scope qualifiers (e.g., ::i); class qualifiers may include template
references (e.g., A<int>::i); and the name that follows these qualifiers
may be more complicated than a simple identifier: it may also be an operator
name (e.g., operator+), a conversion function name (e.g., operator
int), a (C++11) literal operator name (e.g., operator ""_km), or a
destructor name (e.g., ~A). An arbitrarily complex name is referred to as
a “generalized identifier”. The process of collecting the tokens of a
generalized identifier and replacing them with a single tok_identifier
pseudo-token is referred to as “coalescing” the generalized identifier. A
hierarchy of routines is provided that scan generalized identifiers:
is_generalized_identifier_startdetermines whether the current token begins a generalized identifier and, if so, scans the tokens that comprise the generalized identifier and records its description inlocator_for_curr_id.coalesce_and_lookup_qualified_namecallsis_generalized_identifier_startand, if the generalized identifier includes a qualifier, the name after the qualifier is looked up in the scope specified by the qualifier. Names without qualifiers are not looked up. [4]coalesce_and_lookup_generalized_identifieris used to scan and look up a generalized identifier. It callscoalesce_and_lookup_qualified_nameand looks up the name if it does not include a qualifier. Names that include qualifiers will have already been looked up bycoalesce_and_lookup_qualified_name.
The arguments to these routines include options used to specify the kinds of
names that may be accepted and, for the coalesce-and-look-up routines, a lookup
mode that specifies the type of name lookup to be performed. The options
parameter specifies a set of “generalized identifier” (GID) flags. The GID
flags specify several different types of information:
- The kinds of identifiers that should be recognized. For example,
~Amay be recognized as a destructor, or the tilde may be an operator. - The kinds of identifiers that are allowed. For example, are qualified names allowed in this context? Are operator names allowed? The difference between constructs that are “recognized” and those that are “allowed” is that constructs that are not recognized (i.e.,
~D) are not considered identifiers byis_generalized_identifier_startwhile constructs that are not allowed are considered identifiers but cause errors to be issued indicating that the particular type of identifier is not permitted in the current context.
7.3.1. Recognizing Generalized Identifiers#
is_generalized_identifier_start is the lowest-level generalized identifier
routine and does most of the work. As the name suggests,
is_generalized_identifier_start is used to determine whether the current
token is the beginning of a generalized identifier.
is_generalized_identifier_start is actually a macro that first checks
whether the current identifier has already been coalesced. If it has, the
macro immediately returns TRUE; otherwise it calls
f_is_generalized_identifier_start, which performs the rest of the
processing described here.
For these examples is_generalized_identifier_start returns TRUE and sets
the current token to tok_identifier:
i::iX::i::X::i::A::B::iA<int>::iA::operator =A::operator intA::operator ""_km // When user-defined literals are recognizedA::~ANS::iNS::A::ioperator =operator intoperator ""_km // When user-defined literals are recognized~A // When destructors are recognizedA<int>NS::A<int>~int // When nonclass vacuous destructors are recognizedint::~int // When nonclass vacuous destructors are recognizedA::anything else except* // Error
For these examples FALSE is returned and current token is set to
tok_ptr_to_member:
X::*A<int>::* // Template reference will be coalesced
For these examples FALSE is returned and the current token is not changed:
::new::delete~A // When destructors are not recognizedanything else
Determining whether or not the current token is the start of an identifier may
involve scanning all of the tokens that comprise the identifier. Consequently,
all of the tokens that make up an identifier are scanned the first time that
is_generalized_identifer_start is called for a given identifier. The
information needed to describe the generalized identifier is recorded in the
symbol locator, and on subsequent calls the information from the locator is
used to return the appropriate value to the caller. When a sequence of tokens
is coalesced the original tokens are replaced with a single tok_identifier
or tok_ptr_to_member. The resulting token and related state information in
the locator is referred to as a “pseudo-token”. When
is_generalized_identifier_start determines that the thing being scanned is
an identifier, the current token is set to tok_identifier, the locator is
updated with a description of the identifier, and the return value of the
function is TRUE. The fields in the locator that are set are
is_qualified_nameis_global_qualified_nameis_file_scope_qualified_nameis_operator_nameis_destructor_nameis_udl_operator_nameis_class_memberparenthas_been_coalescedis_vacuous_destructor_referenceis_nonclass_destructor
When coalescing a qualified name such as A::xxx the result is a locator
whose symbol header points to the name xxx, and whose parent points to
the A. If A is a class, is_class_member is TRUE and
parent.class_type points to the type of A. If A is a namespace,
is_class_member is FALSE and parent.namespace_ptr points to the
namespace entry for A. The final identifier (xxx) is not looked up.
Instead, the information is preserved in the locator so that, if necessary,
xxx can be looked up by coalesce_and_lookup_qualified_name.
is_generalized_identifier_start coalesces template class references of the
form A::i, [5] A<T>, and A<T>::i (but not simply A) by
calling coalesce_template_id. These references must be coalesced to
determine whether the template reference is the beginning of a qualified name
(e.g., A<int, float>::i cannot be determined to be a qualified name until
the :: is seen.) The result of a template class reference is an incomplete
instantiation. A complete instantiation is only generated when it is required
for a name to be looked up. A complete instantiation will be generated for
references like A<T>::X::Y, because X must be looked up in A<T>. A
complete instantiation will not be generated (at least not here) for references
like A<T>::X. The complete instantiation is delayed until X needs to
be looked up (typically in coalesce_and_lookup_qualified_name.) A template
argument list may also follow a function template name or an overloaded
function name in which one of the members of the overload set is a function
template. This is known as “explicit specification of function template
arguments”. Unlike template class references, a template function reference
cannot be resolved to a given template instance when the template argument list
is scanned (because function templates can be overloaded). As a result, when a
function template argument list is scanned, a pointer to the list is retained
in the symbol locator, but the reference is not otherwise resolved at this
point.
The pointer operator used when declaring a pointer-to-member looks a lot like a
qualified name; the only difference is that the final identifier is replaced
with an asterisk (*). For example, a pointer-to-member operator may look
like A<int>::B::C::*. Because of this similarity pointer-to-member
operators are also coalesced by is_generalized_identifier_start even though
they are not considered identifiers. When a pointer-to-member operator is
coalesced the current token is set to tok_ptr_to_member, the
parent.class_type field of the locator points to the type specified in the
operator (e.g., A<int>::A::B::C in the previous example), and the return
value of the function is FALSE.
The operators ::new and ::delete look like qualified names but are
actually operators. When called for these operators,
is_generalized_identifier_start leaves the current token unchanged and
returns FALSE.
C++ allows a destructor to be called for classes that have no destructor and for nonclass types. We refer to this kind of destructor call as a “vacuous destructor” reference. A GID option is used to indicate whether a vacuous destructor should be recognized and, if so, whether it must be a nonclass destructor. Vacuous destructor references, which are only allowed in field selection operations, differ from other generalized identifiers in several important ways:
- the qualifier, if present, and the identifier that follows the tilde may be a nonclass type.
- the nonclass type found may be a keyword that names a basic type (e.g.,
int::~int).
When a vacuous destructor reference is coalesced, the parent.class_type
points to the type specified by the name after the ::~ because it must be
possible to record the type of the vacuous destructor reference for cases that
do not include a qualifier.
7.3.2. Coalesce-and-Look-Up Routines#
The coalesce-and-look-up routines are:
coalesce_and_lookup_qualified_namecoalesce_and_lookup_generalized_identifier
After is_generalized_identifier_start has determined that the current token
does in fact begin an identifier the coalesce-and-look-up routines are used to
do the appropriate kind of name lookup. The difference between the two
routines is that coalesce_and_lookup_qualified_name will only perform a
lookup of the final identifier of a qualified name while both qualified names
and names without qualifiers will be handled by
coalesce_and_lookup_generalized_identifier.
One of the arguments passed to these routines is an “identifier lookup mode”,
an enumeration that specifies the kind of lookup to be performed. The lookup
mode is translated into a set of “identifier lookup” (IDL) flags that are
passed to the symbol table lookup routines. An error is issued by
coalesce_and_lookup_qualified_name if the name lookup fails, except in
ilm_tentative_type mode. [6] In all other respects the various
identifier lookup modes are identical as far as the coalesce and lookup
routines are concerned. coalesce_and_lookup_generalized_identifier does
not issue an error if the lookup of an unqualified name fails.
7.3.3. Generalized Identifier Errors#
check_for_generalized_identifier_errors is called by the
coalesce-and-look-up routines and also by is_generalized_identifier_start
to ensure that the identifier that has been scanned meets the constraints
specified by the GID “allowed” flags. The error tests are done by comparing
the current GID options with the information in the locator. When one
generalized identifier routine calls another, it masks off the error flags so
that the errors are diagnosed only once by the highest-level routine.
check_for_template_declarator_errors is called by the coalesce-and-look-up
routines to verify that the declarator in a template declaration does actually
refer to a template entity that is permitted in a template declaration. These
checks are done to provide diagnostics that are more helpful than those that
would result if the identifier were simply returned to declarator.
7.3.4. Template References#
coalesce_template_id coalesces references to class and function templates.
The kind of reference to be coalesced is determined based on a symbol that is
provided. If the symbol refers to a function template or an overload set
containing one or more function templates,
coalesce_template_function_reference is called, otherwise
coalesce_template_class_reference is called to handle class template
references or to diagnose the improper use of a template argument list.
7.3.4.1. Template Class References#
coalesce_template_class_reference coalesces template class references, such
as A<int, 1>, into identifiers for which the specific_symbol field of
the locator points to a specific instance of the template. A class template
name is usually, but not always, followed by a template argument list. The
argument list may be omitted in declarations of the class template,
template <class T> class A;
template <class T> class A { /* ... */ };
in which case the name refers to the class template (not to one of its instances), and may also be omitted inside the definition of a class template or in the definition of a member function or static data member of the class template,
template <class T> class A {
A* a; // Equivalent to A<T>* a
int A::* pma; // Equivalent to A<T>::* pma
};
in which case the name refers to the current instance of the class template
(e.g., A<T>).
There are several contexts in which template class references must be
coalesced. Most calls of coalesce_template_class_reference are from
is_generalized_identifier_start and include a template argument list or are
part of a qualified name. coalesce_template_class_reference is also called
by coalesce_and_lookup_generalized_identifier, which scans simple class
template names such as A, and by scan_tag_name, which scans specific
definitions of a template classes such as class A<int> {}. Like the
generalized identifier routines, coalesce_template_class_reference accepts
a set of generalized identifier options as one of its arguments. The only
value actually used is GID_TEMPLATE_ARGS_OPTIONAL which specifies that the
template argument list may be omitted for a reference that is not within the
class template definition. As described above, the argument list may always be
omitted within the class template definition.
The template argument list, if present, is scanned by calling either
scan_template_argument_list or scan_unknown_template_arg_list:
scan_unknown_template_arg_listis used when scanning the template argument list associated with a “nonreal” template – a class template that is a member of a proxy or nonreal class. A nonreal template is created in contexts in which a member class template of a template parameter dependent class is used. For example:template <class T> struct A { typename T::X<int> tx; };
In such cases, it is assumed thatT::Xis a template because there are no cases in which a less than sign could immediately follow a type name, so the “<” must indicate the beginning of a template argument list. When scanning a nonreal template argument list there is no information available about the number of template parameters, their types, etc. All of this must be inferred by scanning the actual arguments. Consequently, no attempt is made to determine whether the number of arguments, the types, etc. are correct. The disambiguation routines are used to determine whether a given argument is a type or nontype (see below for the use of this function in scanning function template references.)scan_template_argument_listis used when scanning the template arguments of “real” templates. Default values are provided for any omitted arguments. If the type of a nontype parameter depends on another template parameter, the type of the parameter is generated by rescanning a token cache constructed when the template was defined. Similarly, if the type of the default value, or the default value itself, involves a template parameter, the default value is also generated by rescanning a token cache constructed when the template was defined.
Both routines call scan_template_argument_constant_expression for nontype
arguments and type_name for type arguments. find_template_class is
called, once a complete template argument list has been created, to find the
type that corresponds to the particular set of arguments. If no such type
exists, find_template_class creates one. In either case the symbol
associated with the type is returned.
7.3.4.2. Function Template References#
coalesce_function_template_reference coalesces references to functions that
employ an explicit template argument list. Unlike class templates, function
templates may be overloaded; therefore an explicit template argument list does
not necessarily identify a specific template instance. The function type of
the instance is also needed to determine the actual instance that is being
referenced, but is not available when the template argument list is scanned, so
the argument list must be retained until some later point when the function
type is also available.
scan_unknown_template_arg_list is used to scan explicit template argument
lists because, in general, there is no template parameter list that can be used
to determine the kind of argument expected (type vs. nontype), or the type of
the argument (for nontype arguments). As with nonreal template argument lists,
the disambiguation routines are used to determine whether a given argument is a
type or nontype. type_name is used to scan arguments that are types.
scan_nontype_template_argument is used to scan nontype arguments. Function
template argument lists differ from those of class templates in that, for
function templates, the corresponding template parameter list is not available
when the template argument list is scanned. The presence of the template
parameter list for a class template makes it possible to convert nontype
arguments to the type of the corresponding template parameter (or to a special
nonreal type in the case of a nonreal argument list). For function templates,
such conversions cannot be done until the specific function template to be used
has been identified. Consequently, when a nontype argument is scanned it must
be retained in a form that permits the conversion to be done later.
scan_nontype_template_argument returns a pointer to an_arg_operand, a
data structure used by the expression scanning routines to represent an
expression. The constant_is_an_arg_operand field in the template argument
is set to indicate that the nontype argument is represented in this form.
7.3.4.3. Variable Template References#
coalesce_variable_template_reference coalesces references to variable
templates. scan_template_argument_list is used to scan the template
arguments. find_template_variable is used to produce a variable symbol for
the instantiation of the given variable template using the specified template
arguments.
7.4. Conversion of Literal Constants#
The source file literals.c contains the code to handle conversion of
literal constants to internal form. The external form is a string of
characters, a token, as accumulated by get_token and its subroutines; the
internal form is a structure of type a_constant, representing an integer,
fixed-point, float, character, or string constant.
conv_integer_literal converts decimal, octal, and hexadecimal integer
constants. Aside from computing the value of the constant, it also recognizes
the trailing L and U suffixes, and determines the type of the overall
constant (on the basis of its value and the suffixes). For pcc mode, the
rules for determining the type of a constant are modified slightly.
conv_fixed_point_literal converts fixed-point constants (an Embedded C
feature), calling fxp_string_to_fixed_point or
fxp_hex_string_to_fixed_point in fixed_pt.c to actually do the
conversion. It also recognizes the various suffixes (k, r, …), and
sets the type of the result accordingly.
conv_float_literal converts float constants, calling fp_string_to_float
or fp_hex_string_to_float_point in float_pt.c to actually do the
conversion. It also recognizes the L and F suffixes, and sets the type
of the result accordingly.
conv_char_literal converts character constants to an integer value. It
uses conv_single_char to fetch and convert each character. If there is
more than one character, they are assembled into one integer (with ordering
controlled by TARG_CHAR_CONSTANT_FIRST_CHAR_MOST_SIGNIFICANT).
conv_string_literal converts string literals. It receives a count of
effective characters in the string from its caller (as computed by
accum_quoted_string), so it can dynamically allocate the space for the
string before scanning it. It also uses conv_single_char to convert each
character of the string.
conv_single_char converts one character to an integer value. This includes
simple characters, escapes like \n, and octal and hexadecimal escapes. At
the moment, this routine assumes that the host and target character sets are
the same.
concat_string_literals concatenates several string literals (in constant
form in a token cache) into one string. This is used for the concatenation of
adjacent string literals.
7.5. Preprocessing Directives#
The source file preproc.c contains the code to handle preprocessing
directives (except for #define, which is handled in macro.c).
Associated declarations are in preproc.h.
The routine cpp_driver is the top-level routine when the front end is being
run as a C preprocessor, typically to produce a preprocessed output file. It
calls get_token repeatedly until end of source, with
do_preprocessing_only set to TRUE to indicate that preprocessing only –
and not compilation – is being done.
Normally, when preprocessing output is required, generate_pp_output is TRUE
as well. But cpp_driver is also called when preprocessing is being done to
generate a list of #include files or a list of makefile dependencies.
In those cases, generate_pp_output is FALSE, and either
list_included_files or list_makefile_dependencies is TRUE.
pp_directive is called (by get_token) when a # is encountered at
the beginning of a source line. It sets in_preprocessing_directive,
fetches the directive identifier and identifies it by calling
identify_dir_keyword (which uses a simple linear search), calls the
appropriate routine to process the directive (which may call
ignore_harmless_trailing_comment to ignore extra text at the end of the
directive), calls end_of_directive_processing (to check for any extra text
that is not harmless), and then resets in_preprocessing_directive. If a
syntax error is detected in a directive, some_error_in_curr_directive is
set; this suppresses the check for extra text at the end of the directive.
When the text of a preprocessing directive is not to be written to the
preprocessing output, do_not_put_curr_line_in_pp_output is set – initially
in pp_directive and on each subsequent line, because of multi-line
comments, in read_logical_source_line; the flag is checked by the output
routine, gen_pp_output_for_curr_line. In the rare case that the text of a
preprocessing directive should be passed to output (because some C compiler
should see the directive to handle it), pass_pp_directive_to_output is set
to TRUE while the directive is scanned.
Each preprocessing directive has an associated processing routine (for example,
proc_if handles the #if directive).
proc_pragma processes the #pragma directive. This is an area that
would be configured for specific versions. Unrecognized pragmas must be
ignored (no error should be generated). When preprocessing only, #pragma
directives are passed to output (see pass_directive_to_output).
proc_ident processes the #ident directive, although the “processing” it
does is to ignore it. When preprocessing only, #ident directives are
passed to output (see pass_directive_to_output).
proc_error passes the rest of the directive to an error reporting routine,
as a catastrophic error, which terminates the compilation.
proc_line scans a line number, treating it as a digit-sequence rather than
an integer constant; optionally, it also scans a file name, calling
copy_header_name to extract the file name from the header name token. It
puts this information into the input stack and calls routines in il.c to
remember the proper correspondence between source sequence numbers and file
name/line number pairs.
proc_include scans the file name (as a header name in one of two forms) and
then calls push_input_stack to request reading from that include file. It
calls copy_header_name to extract the file name from the header name token.
proc_undef scans an identifier, looks it up, and removes any macro
definition under that name. It will not remove predefined macros. (A
predefined macro has a source position with a sequence number of zero, and a
column number of SP_COL_PREDEFINED_MACRO.)
#if and related directives depend on the pp_if_stack, which is a stack
of entries describing active #if constructs. The current depth of the
stack is indicated by pp_if_stack_depth. Each #if that has not yet
been closed by a corresponding #endif appears on the stack. The stack
entry indicates whether or not the #else for the #if has already
appeared, and has the source position of the #if for use as the error
position if the corresponding #endif never appears. In ANSI C or C++ mode,
each #if is required to be ended in the file in which it was begun: to
implement this, the variable base_pp_if_stack_depth contains the value of
pp_if_stack_depth as of the beginning of the current source file. The part
of the stack between base_pp_if_stack_depth and pp_if_stack_depth is
considered the only active part of the stack at any given moment. #endifs encountered must be associated with #ifs in the active part of the
stack, and all of the #ifs in the active part of the stack must be closed
by the end of the source file (this is checked by
verify_that_all_pp_ifs_were_closed, which is called from
pop_input_stack). In pcc mode, the #ifs do not have to match up
within each file; they only have to come out right at the end of the
compilation.
The various if-directives are handled by proc_if, proc_ifdef (which
also handles #ifndef), proc_else, proc_elif, and proc_endif.
proc_if and proc_elif call scan_if_expr, which calls
scan_pp_expression (in expr.c) to scan the if-expression (there are
restrictions on what can appear in such an expression). All of the
if-processors call perform_if once they know whether or not the #if
should be executed; perform_if returns if the condition is TRUE, and calls
skip_to_endif if the condition is FALSE. The latter reads lines looking
for preprocessing directives, counts nesting of #ifs, and looks for an
#else or #endif that matches the current if.
A new preprocessing directive can be added by:
- Adding an enumeration constant to
a_pp_directive_kind; - adding a test for the directive keyword in
identify_dir_keyword; - adding a switch clause for the directive in
pp_directive; and - adding a routine
proc_directive to process the directive. Note that the routine must take all the tokens in the directive (up to thetok_newline) before returning.
7.5.1. AT&T Preprocessing Extensions#
When ATT_PREPROCESSING_EXTENSIONS_ALLOWED is defined some of the AT&T
System V release 4 extensions to the preprocessor are supported. In that case,
the #assert and #unassert directives are handled by proc_assert and
proc_unassert. The data structure built for assert predicates is a simple
one: assert_predicates points to a list of an_assert_predicate entries,
each of which gives a predicate name and a list of an_assert_value entries
giving the value or values for the predicate. Entry and lookup are done with a
simple linear search. The token sequences involved are processed into
character strings (containing the token characters with a blank after each
token, and ended with a null character). Such character strings then represent
the token sequences, and can be saved and compared.
scan_assert_predicate_reference is called from get_token on detection
of a # inside of an #if expression; it scans the reference and returns
1 or 0 depending on whether or not the predicate is defined with the indicated
token sequence. enter_assert_predicate can be called from fe_init.c to
enter predefined predicate names.
7.5.2. C++/CLI Assembly Import#
The front supports the C++/CLI #using directive to import the metadata of
an assembly. This is done in function proc_using. In C++/CLI mode, a
“core library assembly” (usually, mscorlib.dll) is imported before input
source is processed, as well as any assemblies specified with the
--preusing command-line option: Function process_preusings handles that
part (see also Declarations from C++/CLI Assembly Metadata and C++/CLI Initialization).
7.6. Macro Processing#
macro.c contains the code to handle the #define preprocessing directive
and invocations of macros. macro.h contains associated declarations.
7.6.1. Macro Definition Data Structure#
When a macro is defined, the symbol created for it points to an entry of type
a_macro_def (defined in symbol_tbl.h). The macro’s parameter list is
represented by a linked list of entries of type a_macro_param; these give
the names of the parameters (needed only to verify that, on a benign
redefinition of the macro, the parameter names are the same). The replacement
text of the macro is a sequence of sections, each of which indicates text that
should be placed in the macro expansion at that point. Each section begins
with a byte that identifies the kind of section, as follows:
rt_null (0) |
marks the end of the replacement text.
|
rt_text |
indicates raw text that is part of the macro replacement text. The
next three bytes contain the count of characters following.
|
rt_raw_argumentrt_right_raw_argument |
indicate that the raw (non-macro-expanded) text of an argument is to
be inserted at this point. This results from use of the
##
operator, or from parameters scanned in pcc mode. The next three
bytes contain a number indicating the argument to be used (1 is the
first argument, 2 the second, etc.). rt_raw_argument is used for
a parameter to the left of ## and for all parameters in pcc mode,
and rt_right_raw_argument is used for a parameter to the right of
##. |
rt_stringized_raw_argument |
indicates that the “stringized” text of an argument is to be inserted
at this point (this results from use of the
# operator). The next
three bytes contain a number indicating the argument to be used. |
rt_argument |
indicates that the macro-expanded text of an argument is to be
inserted at this point. The next three bytes contain a number
indicating the argument to be used.
|
The ## (paste) operator has no direct representation in the replacement
text; it just results in the two entities being placed next to each other. For
example,
#define x(a,b) a ## b
results in rt_raw_argument(1) followed by rt_right_raw_argument(2).
White space in the replacement text is standardized. Any white space preceding
or following the replacement text is removed, as is any white space preceding
or following a ##, and any white space following a #. All other white
space is replaced by a single space.
Each token in the raw text (except the last, and any token preceding a ##)
is followed by an LE_END_OF_TOKEN lexical escape sequence. The raw text
immediately following insertions of arguments (except immediately following a
##) also starts with an end-of-token marker, to terminate the last token of
the inserted argument text. The end-of-token markers ensure that the tokens
will be tokenized the same way each time. For example, one problem comes up in
#define z(a) ..##a
The value of z(.) should be three “.” tokens, not the one token
“...”. End-of-token markers are not used when in pcc mode, since cpp
does not operate on a token level.
Similar standardization of white space and insertion of end-of-token markers is done when the raw form of macro arguments is accumulated. This standardization (of replacement text and macro argument text) makes it easier to check for benign redefinitions and to generate the stringized form of arguments.
7.6.2. The Macro Buffer#
The global variable macro_buffer points to a large character array used as
a work area during macro processing. During macro definition, it holds the
replacement text. During macro invocation, it holds the expansions of all
macro invocations in the current line. It is dynamically allocated, and can be
expanded as necessary by calling expand_macro_buffer.
expand_macro_buffer will, in turn, call
adjust_curr_source_line_structure_after_realloc to walk through the data
structure associated with curr_source_line and change pointers to
macro_buffer to reference addresses within the new allocation for
macro_buffer. adjust_curr_source_line_structure_after_realloc
automatically finds such pointers in global variables and the public data
structure, but it has no way of knowing about local pointer variables.
Therefore, routines that have such variables register them by calling
register_pointer_variable, so that they will be found and adjusted if
reallocation is done.
In order to reduce memory consumption when processing very complex and lengthy
macro expansions, the reallocation process for macro_buffer differs from
that of the other extensible buffers. While the other buffers are simply
reallocated, relying on the C runtime to copy the complete old contents to the
new allocation, expand_macro_buffer compacts the contents of
macro_buffer by explicitly allocating a new buffer, copying data to the new
allocation only if it might still be in use, and then freeing the old buffer.
First, each existing source line modification (see
Modifications to the Current Source Line) is processed to copy its inserted text.
Within that text, only the ATTENTION_MARKER is copied from each portion
that is marked for deletion by another source line modification. (The
ATTENTION_MARKERs must be copied in order to allow scanning for inert
macro names, as described in Macro Invocations.) As a result, the
original text of already-expanded macro invocations and the expanded text of
macro arguments will not occupy space in the new macro_buffer allocation.
Second, text that is in the process of being constructed at the end of the
macro_buffer is copied to the end of the data in the new allocation. Any
routines that build data structures incrementally, so that reallocation can
occur before the structure is complete (e.g., macro replacement text in
proc_define), must identify the beginning of such data using
begin_macro_buffer_region and its completion with
release_macro_buffer_region.
7.6.3. Macro Definitions#
Macro definitions (#define directive) are handled by proc_define. It
scans the macro identifier, and looks it up in the symbol table (predefined
macros may not be redefined, and normal macros may be redefined only if the new
parameters and text match the existing parameters and text – a “benign”
redefinition). If a “(” follows immediately (with no intervening white
space), the macro is function-like; its parameter names are accumulated as a
list. Otherwise, the macro is object-like. proc_define then scans the
replacement text and builds the replacement text string in macro_buffer.
It calls mdefn_get_token to fetch tokens of the body and remember whether
or not any white space preceded each. White space is standardized,
end-of-token markers are inserted (if in ANSI C or C++ mode), parameter names
are recognized, and # and ## are handled as described above.
If in pcc mode, the opening and closing quoting characters of character
constants and string literals are scanned as single-character tokens, and the
insides of the quoted strings are tokenized. This results in recognition (and
later replacement) of parameter names within strings, as cpp does it.
Invalid tokens are passed textually to the replacement text (this is necessary
always, not just for this case), and skipped white space is kept in its
original form, so the contents of the string will be accurately rendered in the
final expansion of the macro.
After the replacement text is accumulated, it is copied to dynamically-allocated storage, and the macro definition is completed and attached to the symbol. For a redefinition, the new parameter list and replacement text are first compared textually against the old. If they match, the new definition is thrown away.
7.6.4. Macro Invocations#
macro_invocation handles the scanning and replacement of macro invocations.
It is called by get_token when an identifier that is a macro name is
scanned.
The first thing that macro_invocation checks is whether or not the macro
name is “inert”. The standard says that when a macro’s name is generated in
its own expansion (directly or indirectly), it is forever after protected
against further expansion (it is inert). This is implemented by the
assoc_macro field in the a_source_line_modif entry. A macro invocation
is replaced by the macro’s expansion by creating a modification entry that
indicates the textual changes to be made to the source line or to a previous
modification. The assoc_macro field records the origin of any given
change, by pointing to the definition of the macro that generated the change.
Given a pointer to the name of a macro at the beginning of a possible macro
invocation, one can therefore determine whether it is in the main source line
or in a modification (or nested in several), and from that one can see whether
the macro name appears as part of its own expansion. If so, the macro name is
inert, and is left as an identifier. (Actually, it is replaced by an
LE_INERT_MACRO lexical escape followed by the macro name. Once a macro
name is preceded by an inert-macro escape, it is considered inert every time it
is scanned, regardless of the context.) In pcc mode, macros are not considered
to be inert, and recursion is checked for by counting the number of nested
invocations of a given macro.
If the macro name is not inert, processing for expansion continues.
delete_source_from_loc is set to point to the beginning of the macro
identifier. This begins a “hanging delete” which will be closed when the end
of the macro invocation is found. At that point, the text of the hanging
delete will be replaced by the macro expansion, this replacement being
indicated by a source line modification entry.
If the macro is object-like, no parameters need be scanned. However, there are
several special cases in object-like macros: For __FILE__ and __LINE__,
the proper expansion string is developed from the current position information
in the input stack; for defined, which is an operator allowed in #if
expressions but handled as a pseudo-macro, the routine
scan_defined_operator is called. It scans both forms of the defined
operator, and returns 0 or 1. Alternatively, if the identifier
defined is not followed by appropriate tokens, it is left as an identifier
(see the discussion of check_for_following_parenthesis below).
If the macro is function-like, check_for_following_parenthesis is called to
scan forward looking for a parenthesis that begins the argument list. If the
next token is not a parenthesis, then the macro name is left as an identifier.
However, this is complicated, because in skipping white space while looking for
a left parenthesis, we may have gone onto a new source line. We have to delete
the identifier before leaving the old source line, on the assumption that we
are scanning a macro invocation. If it turns out that the identifier is not
the start of a macro invocation, the identifier must be reinserted into the new
source line.
Note the field sequence_id in a_source_line_modif; it is an integer
sequence number indicating the sequence in which source line modifications were
added. This is useful in cases where modifications must be undone (as in
identifier un-deletion when a macro identifier is followed by something other
than a left parenthesis on the same line); one can find all entries with
sequence numbers greater than a certain number, and remove them.
If a left parenthesis is found, the argument list is scanned, and this
requires a new data structure. An entry of type a_macro_arg contains
the text of a single argument value. That text exists in two forms – as
“raw” text (not macro-expanded), and as “expanded” text. The
a_macro_arg entries are dynamically allocated, and freed when no longer
needed (see alloc_macro_arg, free_macro_arg, and the variable
avail_macro_args). The text of macro arguments is not kept in
macro_buffer.
Each macro argument is initially scanned in raw form. arg_get_token is
called, with macro expansion turned off, and the characters of each token, with
standardized white space and end-of-token markers, are placed in the array
pointed to by raw_text in the a_macro_arg entry (a
dynamically-allocated array, which can be reallocated as needed to expand it).
The scan stops when a comma or right parenthesis is found. A variable is used
to count the nesting of parentheses, so that non-zero-level commas and right
parentheses will not stop the scan.
After the raw form of the argument has been accumulated, the raw text is
temporarily re-inserted in the source line and is re-scanned with macro
expansion on. The standard says that this must be done as if this text forms
the rest of the source file (i.e., the macro expansion is done in isolation
from the text following the raw argument value). This is implemented by having
the source modification entry that re-inserts the raw text have the field
is_isolated_text TRUE. This will cause get_token and
skip_white_space, upon reaching the LE_END_OF_INSERTION lexical escape
sequence at the end of the raw text, to return tok_end_of_source instead of
continuing into the surrounding text. macro_invocation scans the entire
raw argument again, saving the text of the tokens scanned in the
expanded_text array in the macro argument entry, and stopping when it gets
tok_end_of_source. If any modifications of the raw text were done because
of macro expansions (this can be determined by sequence id), they are removed.
This restores the raw text to its original state. In pcc mode, and when the
definition of the macro does not use the parameter in a context that requires
expansion,only the raw form of each argument is needed, so the rescan is not
done.
In cases where the macro argument is the result of a previous macro expansion,
the rescan begins in the text inserted by the macro expansion modifications
rather than in the raw_text buffer. That ensures that the inertness test
can be done properly on macro names encountered in the rescanned source. Once
the rescan reaches the end of the insertions and would continue into the
primary source line, the input stream is switched to the tail of the
raw_text buffer. (raw_text is used instead of the primary source line
because macro arguments can be continued to subsequent source lines, and once
the next line is read the only copy of the raw argument text is in the
raw_text array; the previous source line is gone.)
This process is repeated for each argument. After all arguments are scanned
(or immediately for object-like macros), expansion is done. The length of the
expansion is first determined, and then the textual change is made, by building
up text in macro_buffer and placing a source line modification to delete
the characters of the macro invocation and insert the characters of the
expansion.
For the special cases of predefined macros, the replacement text is already available as a simple null-terminated string. For these cases, determining the length and building the replacement string are easily done.
For the other, more normal, cases, the length is determined and the expansion
string is built up by interpreting the sections of the replacement text in
order. For unexpanded uses of parameters, the raw_text is copied. For
macro-expanded uses of parameters, the expanded_text is copied. For
stringized parameters, stringized_arg is called; it copies the
raw_text, inserting the surrounding quotes and the additional backslashes
to protect quotes and backslashes in the string.
The source line modification is then inserted to perform the macro expansion.
When, in pcc mode, the expansion is for a top-level macro,
expand_top_level_pcc_macro is called to expand any macro calls within the
body of the macro. A complete string for the expanded version is constructed
in an auxiliary buffer (aux_buffer_for_pcc_macros) and then copied back
into macro_buffer, replacing the original expansion and any source line
modifications associated with it. Getting a fully-expanded string is necessary
to get pcc-like token pasting behavior across internal macro expansion
boundaries. A variant of this processing is also done in Microsoft mode.
macro_invocation then returns. It indicates to get_token whether or
not the inserted text should be rescanned. If macro_invocation already
knows what the next token should be, it returns it directly (*rescan =
FALSE), thus saving get_token the time and trouble of scanning the token
text again.
A utility debug routine print_markered_text is available for use in
printing out text strings that contain end-of-token markers and/or transitions
into or out of source modifications.
7.6.5. Macro Position Tracking#
When FULLY_RESOLVED_MACRO_POSITIONS is TRUE, the front end keeps track of
the original location of text that is copied into macro expansions – i.e.,
either in a #define line or in an argument to a top-level macro invocation.
This “original position” information is reflected in the orig_seq and
orig_column fields of a_source_position.
The principal data structure used in tracking original position information
through the various manipulations performed during macro definition and
invocation is a_macro_text_map, which defines an extensible array of
a_macro_text_map_entry elements (both defined in lexical.h). Each of
the repositories of text associated with macro definition and invocation –
macro definitions, the raw and expanded forms of macro arguments, and source
line modifications – has an associated macro text map.
A macro text map provides an ordered sequence of correspondences between
offsets in the associated text buffer and the original source locations from
which that text was copied. Each map entry defines a region of the associated
text buffer, beginning at offset start_of_region and extending up to but
not including the start_of_region of the next entry. The text at offset
start_of_region originally came from the location given by
corresponding_source_pos, with a one-to-one correspondence of increasing
offsets and source column positions throughout the entry’s buffer region.
Because the entries are ordered by buffer offset, the original source position
for a given offset can be efficiently found using a binary search (the
requisite bsearch predicate is
compare_macro_text_map_entry_with_offset, defined in lexical.c).
It will be helpful to trace the overall flow of information through the
definition and invocation of a macro before covering the details of how it is
accomplished. When a #define is processed, proc_define builds a text
map by incrementally recording the source position and offset in the
replacement text for each token of the definition that will be copied to the
macro expansion. A similar process occurs as each macro argument is copied to
the corresponding raw and expanded text buffers. During the interpretation of
the replacement text by macro_invocation, as each section of text is copied
from the replacement text or macro argument, the corresponding entries from the
associated macro text map are copied, adjusting the offsets to reflect the
location of the text in the macro expansion (but leaving the original source
positions). Finally, when conv_line_loc_to_source_pos is called to compute
a source position for a given character location in the expanded text, the
macro text map in the source line modification is used to determine the
corresponding original source position.
There are two principal patterns for adding entries to a text map. The simpler
pattern is when a sequence of map entries is copied from one map to another,
adjusting the offsets from being relative to the start of the source buffer to
being relative to the start of the target buffer. This technique is used to
create the text map for the source line modification containing the macro
expansion by copying portions of the text maps for the macro definition
replacement text and macro arguments. The function that performs this
operation is clone_macro_text_map_entries, defined in macro.c.
The other pattern occurs when a text buffer is built incrementally by fetching one token at a time and adding its text to a buffer. In this case, the map entries are built to reflect either the source position of the token, if the token is in source text, or a copy of the corresponding text map entry, if the token comes from a source line modification (e.g., a macro argument in a macro invocation appearing in the expansion of another macro invocation). This technique is used to create the replacement text in a macro definition and for both the raw and expanded text of macro arguments. It is also used when processing a top-level macro invocation in pcc and Microsoft modes.
Facilitating this process are the struct a_text_map_position_tracker and
the functions init_text_map_position_tracker,
add_token_to_macro_text_map, and terminate_macro_text_map (all defined
in macro.c). As each token is added to the text buffer,
add_token_to_macro_text_map is called to determine whether a new entry is
required in the map for the new token: if the token is from the same source
line or source line modification as the preceding token and its offset within
the current buffer region is the same as the column offset from the region’s
corresponding source position, then no new map entry is needed. If a new map
entry is required, either the appropriate map entries from the previous token’s
source line modification are copied, using clone_macro_text_map_entries, or
(for tokens from the current source line) a new entry is added covering the
appropriate region of the current source line for the preceding token(s).
One complication in this processing occurs when macro_buffer is
reallocated, as described earlier. Because a_text_map_position_tracker
objects in use at the time macro_buffer is reallocated might contain the
offset of text in a source line modification’s inserted text that occurs after
a portion of deleted text, and because that deleted text might be compacted as
part of the reallocation process, it is necessary to keep track of all active
position tracker objects so their current offset values can be adjusted if
necessary.
This is done using the variable active_text_map_position_trackers (in
macro.c), which is the top of a stack of all trackers currently in use, and
the field num_active_position_trackers in a_source_line_modif.
Whenever a token scan under control of a text map position tracker enters the
inserted text of a source line modification, its
num_active_position_trackers is incremented, and it is decremented when the
scan leaves that source line modification.
Whenever macro_buffer is reallocated, all the entries in each source line
modification’s text map are examined and each offset adjusted, if necessary, to
compensate for the removal of deleted text. Furthermore, if the modification’s
num_active_position_trackers field is greater than zero, all the trackers
in the active_text_map_position_trackers stack are checked for references
to that source line modification and their offsets adjusted as needed.
The text maps for macro definitions and the raw and expanded text of macro
arguments are completely self-contained: each has its own array of text map
entries and is associated with the containing data structure for the duration
of the compilation. (The array of entries for a macro definition is of fixed
size and is allocated in front-end memory, so that it will be stored in a
precompiled header, while the arrays for macro arguments are extensible.) A
different scheme is used for the text maps in source line modifications,
however. The reason for this decision is that the text map for a source line
modification may well be very large, with thousands or tens of thousands of
entries, while source line modifications themselves are ephemeral, generally no
longer needed once processing of the current logical source line has been
completed. The same reasoning led to the use of macro_buffer for the
inserted text of source line modification, so the memory management for source
line modification text maps is patterned after that of macro_buffer. Just
as each modification’s inserted text is simply a subrange of the text in
macro_buffer, the entries created during macro expansion are built in a
single map, called macro_text_map and defined in macro.c, and each
modification’s text map simply refers to a subrange of them. Initialization
and truncation of macro_text_map is done in parallel with the corresponding
operations on macro_buffer.
7.6.6. Macro Invocation Tree#
Maintaining the macro invocation tree is selected by setting
MACRO_INVOCATION_TREE_IN_IL to TRUE. Each macro invocation is recorded in
an object of type a_macro_invocation_record (defined in il_def.h),
giving the parent invocation, if any; the macro invoked; and its starting
position (as well as its ending position, if EXTRA_SOURCE_POSITIONS_IN_IL
is set to TRUE). (The “parent invocation” refers to the macro invocation in
whose argument list or expansion this macro invocation occurred. For macro
invocations occurring in program text and not in a macro argument list,
parent_macro_index will have the value NO_PARENT_MACRO_INVOCATION.)
Each time macro_invocation is called to expand a macro, it calls
register_macro_invocation to add an entry to the list and obtain the index
of the newly-added entry. This index and the macro invocation stack depth at
that point are stored into the source line modification created to hold the
macro’s expanded text. These values are used by recursive calls of
macro_invocation (to expand macro arguments) and by calls resulting from
rescanning expanded text to maintain the chain of parent invocations when the
macro name comes from a source line modification rather than directly from
source text.
Because variable-length data is awkward to represent in the IL (and extensible
arrays are not possible in memory regions), the macro invocation records are
grouped into fixed-size blocks (a_macro_invocation_record_block, defined in
il_def.h). In the IL, these blocks are organized into a binary tree to
facilitate relatively-efficient random access by index. During front-end
processing, however, they are kept in a doubly-linked list: apart from
diagnostic output, which need not be terribly efficient, there is no need for
random access to the invocation records (and even diagnostic output will most
often refer to records near the end of the list, so long linear searches are
uncommon). It is more efficient to create the binary tree form once, after the
total number of records is known, than to balance the tree repeatedly while it
is being created.
The macro_invocation_record_blocks are reorganized into a binary tree and
the relevant fields of il_header are set at the end of the translation unit
by a call to copy_macro_invocation_tree_to_il (defined in macro.c).
An additional effect of setting MACRO_INVOCATION_TREE_IN_IL is that
a_source_position is extended to include the macro_context field. This
field reflects the macro invocation in whose expansion the source position
resides, i.e., it is logically simply a copy of the invocation_record field
of the containing source line modification. However, the special treatment
given to top-level macros in pcc and Microsoft mode, in which the expansion of
the macro is rescanned for expansion and the result is copied back into the
source line modification corresponding to the top-level macro invocation, means
that the subsidiary source line modifications are no longer available at the
time source positions are being computed.
To avoid this problem, the macro invocation index that will be stored by
macro_invocation into source line modifications is also kept in the
macro_context field of each text map entry created for that invocation.
This information is preserved when the text maps for the subsidiary source line
modifications are cloned into the top-level modification by
expand_top_level_pcc_macro. This additional information allows the
macro_context field in source positions to be set in the same way as the
original position information, by looking up the offset in the text map of the
top-level modification. (This is the reason that it is not permitted to set
MACRO_INVOCATION_TREE_IN_IL to TRUE without also setting
FULLY_RESOLVED_MACRO_POSITIONS to TRUE.)