3. Overview of Internal Structure#
3.1. Function of the Front End#
The front end reads source files written in C or C++, diagnoses any errors in them, and produces an intermediate language tree that represents the source program.
The source files are one primary source file (named on the command line)
and zero or more header files (specified by #include directives). If
COMPILE_MULTIPLE_SOURCE_FILES or COMPILE_MULTIPLE_TRANSLATION_UNITS
is TRUE, multiple source files may be specified on the command line.
The intermediate language tree is an in-memory representation of the source program, from which a back end can generate object code. This intermediate form is high-level (i.e., its constructs correspond closely to C or C++ constructs). Implicit transformations in the source language are made explicit in the intermediate language. Source-correspondence information is preserved so that symbolic debugging information can be generated.
When it translates C programs to intermediate language, the front end uses only the C subset of the intermediate language. When it translates C++ programs to intermediate language, it uses the full intermediate language. However, an optional IL lowering pass can be used translate the C++ intermediate language constructs into C intermediate language constructs (this lowering is not available for C++/CLI and C++/CX).
The front end does not do any optimizations. It does not transform the program in other than trivial ways.
The front end would usually be run from a driver, which invokes the front end and back end in order, and which validates and distributes the command-line options. In such a case, the front end would write the intermediate language information to a file, which would be read by the back end.
The front end incorporates the preprocessing function of the C and C++ languages, and does them on the fly as it parses the source program. The front end can also be run in a preprocessing-only mode, and will generate preprocessing output as a separate preprocessor would. Comments can be kept or removed in this process.
The front end can generate a file containing raw listing information to be used to assemble a compilation listing. The raw listing file contains raw source lines, information on transitions into and out of include files, and diagnostics generated by the front end. A file containing cross-reference information (information on each reference to an identifier in the source program) can also be generated in a different file.
3.2. Guide to the Source Files#
Note that most source files have a .c file suffix, even though they contain
C++ code and must be compiled with a C++ compiler (with the C++11 dialect).
That naming convention is for historical reasons and for ease in delivering
patches against earlier versions of the front end. Symbolic links (e.g., with
a .cpp suffix) can be added if required by a build environment.
The source files are:
attribute.h |
Declarations related to
attribute.c (scanning of attributes). |
basics.h |
Basic declarations for every compilation.
|
builtin_defs.h |
Declarations related to GCC/clang/Microsoft builtin functions. This
header file is generated automatically by a tool that retrieves the
signatures of these builtin functions from various versions of GCC and
clang compilers (the Microsoft builtin function signatures are
manually input to this automated process).
|
builtin_kinds.h |
Builtin function kinds (split from
builtin_defs.h to improve
compilation times). This header file is generated automatically by a
tool that retrieves the signatures of these builtin functions from
various versions of GCC and clang compilers (the Microsoft builtin
function signatures are manually input to this automated process). |
c_gen_be.h |
Declarations related to
c_gen_be.c (C-generating back end). |
cfe_daemon_common.h |
Declarations related to
cfe_daemon.c exposed for utility programs. |
checking.h |
Fundamental declarations associated with assertion checking.
|
class_decl.h |
Declarations related to
class_decl.c (class declarations
scanning). |
cmd_line.h |
Declarations related to
cmd_line.c (command-line parsing). |
const_ints.h |
Declarations related to the manipulation of the internal
representation of integer constants.
|
cp_gen_be.h |
Declarations related to
cp_gen_be.c (C++/C-generating back end). |
debug.h |
Declarations related to
debug.c (debug output). |
decl_inits.h |
Declarations related to
decl_inits.c (declaration initializers
scanning). |
decl_spec.h |
Declarations related to
decl_spec.c (declaration specifiers
scanning). |
declarator.h |
Declarations related to
declarator.c (declarator scanning). |
decls.h |
Declarations related to
decls.c (declarations scanning). |
def_arg.h |
Declarations related to
def_arg.c (default argument scanning). |
defines.h |
A place to put
#defines to configure the front end;
defines.h is included at the beginning of every source file. |
direct_allocator.h |
Declaration and definition of the
Direct_allocator template. |
disambig.h |
Declarations related to
disambig.c (C++ expression-declaration
disambiguation). |
error.h |
Declarations related to
error.c (error reporting). |
err_codes.h |
Definition of the enumeration of error codes. This file is generated
from the
error_msg.txt file. See mk_errinfo. |
err_data.h |
Definition of the enumeration of error codes. This file is generated
from the
error_msg.txt and error_tag.txt files. See
mk_errinfo. |
expr.h |
Declarations related to
expr.c (expression scanning). |
exprutil.h |
Declarations related to
exprutil.c (expression scanning
utilities). Also, the data structures used in expression scanning. |
extasm.h |
Declarations related to
extasm.c (GNU C extended asm statement
scanning). |
fe_init.h |
Declarations related to
fe_init.c (initialization). |
fe_wrapup.h |
Declarations related to
fe_wrapup.c (termination). |
fixed_pt.h |
Declarations related to
fixed_pt.c (manipulation of fixed-point
constants). |
float_pt.h |
Declarations related to
float_pt.c (manipulation of floating-point
constants). |
float_type.h |
Definitions of floating-point conversion routines used when
USE_HOST_FP_CONVERSION_ROUTINES is FALSE. This header file is
included multiple times (once for each distinct floating-point type)
by floating.c. |
floating.h |
Declarations related to
floating.c (software-based conversion of
floating-point values between binary and decimal). |
folding.h |
Declarations related to
folding.c (folding of constant
operations). |
func_def.h |
Declarations related to
func_def.c (processing for function
definitions). |
header_util.h |
General utility components (mostly templates) intended for use in both
headers and compilation units.
|
host_envir.h |
Declarations related to
host_envir.c (definition of the host
environment, including operating system dependencies). Also, many
host configuration switches. |
host_util.h |
Definitions of host environment functions that are used by both the
front end and utility programs such as the prelinker. When the front
end is compiled, these routines are included as part of
host_envir.c. |
ifc_map.h |
Automatically-generated type declarations and definitions for
interacting with the IFC modules format.
|
ifc_map_functions.h |
Automatically-generated function declarations for interacting with the
IFC modules format.
|
ifc_modules_inst.h |
Automatically-generated macro invocations to facilitate instantiations
in
ifc_modules_templ.c. |
ifc_modules_internal.h |
Declarations related only to the internal
implementation files for IFC modules file support.
|
ifc_modules_spec.h |
Automatically-generated macro invocations to facilitate explicit
specializations in
ifc_modules_templ.c. |
ifc_modules.h |
Declarations related to IFC modules files.
|
il.h |
Declarations related to
il.c (intermediate language). |
il_alloc.h |
Declarations related to
il_alloc.c (allocation of intermediate
language entities). |
il_def.h |
The definition of the intermediate language.
|
il_display.h |
Declarations related to
il_display.c (display of intermediate
language in human-readable form). |
il_file.h |
Declarations related to the structure of the intermediate language
file.
|
il_read.h |
Declarations related to
il_read.c (reading the intermediate
language file). |
il_to_str.h |
Declarations related to
il_to_str.c (creating external string-form
representations for certain IL entries). |
il_walk.h |
Declarations related to
il_walk.c (walking the intermediate
language tree). |
il_write.h |
Declarations related to
il_write.c (writing the intermediate
language file). |
inline.h |
Declarations related to
inline.c (minimal inlining for IL
lowering). |
interpret.h |
Declarations related to
interpret.c (IL interpreter for
constexpr functions). |
lang_feat.h |
Declarations of configuration switches that affect the source language
accepted.
|
layout.h |
Declarations related to
layout.c (class object layout). |
lexical.h |
Declarations related to
lexical.c (source input and lexical
scanning). Also, the definition of the data structures related to the
current input line. |
lint.h |
A place to put directives to silence the
lint program. |
literals.h |
Declarations related to
literals.c (conversion of literal
constants). |
lookup.h |
Declarations related to
lookup.c (name lookup). |
lower_c99.h |
Declarations related to
lower_c99.c (C99 IL lowering). |
lower_eh.h |
Declarations related to
lower_eh.c (IL lowering of exceptions). |
lower_il.h |
Declarations related to
lower_il.c (IL lowering). |
lower_init.h |
Declarations related to
lower_init.c (IL lowering of
initializations). |
lower_name.h |
Declarations related to
lower_name.c (name mangling for IL
lowering). |
macro.h |
Declarations related to
macro.c (macro definition and invocation). |
mem_manage.h |
Declarations related to
mem_manage.c (memory management). |
mem_tables.h |
Declarations related to the memory management data structure.
|
modules.h |
Declarations related to the processing of module files.
|
ms_attrib.h |
Declarations related to Microsoft attribute processing.
|
ms_metadata.h |
Declarations related to
ms_metadata.cpp (reading C++/CLI metadata
from assemblies). |
overload.h |
Declarations related to
overload.c (overload resolution for
expression processing). |
pch.h |
Declarations related to
pch.c (precompiled header processing). |
pragma.h |
Declarations related to
pragma.c (processing of #pragma
directives). |
preproc.h |
Declarations related to
preproc.c (preprocessing directives). |
scope_stk.h |
Declarations related to
scope_stk.c (scope stack management). |
src_seq.h |
Declarations related to
src_seq.c (source sequence list
management). |
statements.h |
Declarations related to
statements.c (statement scanning). Also,
the definition of the statement stack. |
symbol_ref.h |
Declarations related to
symbol_ref.c (symbol references). |
symbol_tbl.h |
Declarations related to
symbol_tbl.c (symbol table). Also, the
definitions of the symbol table and the scope stack. |
sys_predef.h |
Declarations related to
sys_predef.c (system-specific predefined
macros). Also declarations related to builtins (and the place to add
user-defined builtins). |
targ_def.h |
Declarations that define default target machine characteristics, and
target configuration switches.
|
target.h |
Declarations of target configuration variables, initialized to default
values defined in
targ_def.h. |
target_cfg.h |
Definitions of target-specific routines to set and dump target
configurations (included once for each target configuration).
|
target_map.h |
Definition of various routines that map target-specific configuration
macros to global variables in some fashion (included multiple times).
|
templates.h |
Declarations related to
templates.c (class and function
templates). |
trans_copy.h |
Declarations related to
trans_copy.c (copying of secondary
translation unit IL to primary IL). |
trans_corresp.h |
Declarations related to
trans_corresp.c (establishment of
correspondences between translation units). |
trans_unit.h |
Declarations related to
trans_unit.c (driver for translation-unit
processing). |
types.h |
Declarations related to
types.c (utility routines dealing with
types). |
util.h |
General utility components. This contains a set of (mostly templated)
utilities that are used in place of those that would normally be found
in a C++ standard library. The declarations are conditionally
enclosed in the
edg namespace (under control of the
USE_EDG_NAMESPACE configuration macro). |
version.h |
Definition of the current version number of the front end.
|
walk_entry.h |
Routines used by
il_walk.c to walk through IL entries. |
attribute.c |
Scanning of attributes.
|
c_gen_be.c |
The C-generating back end (for use in place of a “real” back end;
generates C code which can be compiled by a native target compiler).
|
cp_gen_be.c |
The C++/C-generating back end (for use in source-to-source
transformation).
|
cfe.c |
The main program for the front end.
|
cfe_daemon.c |
The main program for the front end when built as a daemon..
|
class_decl.c |
Class declarations scanning.
|
cmd_line.c |
Command-line processing.
|
const_ints.c |
Manipulation of the internal representation of integer constants.
|
debug.c |
Debug control routines.
|
decl_inits.c |
Declaration initializers scanning.
|
decl_spec.c |
Declaration specifiers scanning.
|
declarator.c |
Declarator scanning.
|
decls.c |
Declarations scanning.
|
def_arg.c |
Default argument caching and rescanning.
|
disambig.c |
C++ expression-declaration disambiguation.
|
error.c |
Error reporting.
|
expr.c |
Expression scanning.
|
exprutil.c |
Utilities for expression scanning.
|
extasm.c |
Scanning of GNU C extended asm statement.
|
fe_init.c |
Overall initialization.
|
fe_wrapup.c |
Overall termination.
|
fixed_pt.c |
Manipulation of fixed-point constants: conversion to and from internal
form and conversions and operations on the internal form.
|
float_pt.c |
Manipulation of floating-point constants: conversion to and from
internal form and conversions and operations on the internal form.
|
floating.c |
Conversion of floating-point values between binary and decimal
formats. Used only when
USE_HOST_FP_CONVERSION_ROUTINES is
FALSE. |
folding.c |
Folding of constant operations.
|
func_def.c |
Processing for function definitions.
|
host_envir.c |
Handling of host-environment dependencies (e.g., file name rules and
opening of files).
|
ifc_map_functions.c |
Automatically-generated function definitions for interacting with the
IFC modules format (that do not fit into one of the more specific IFC
map functions files).
|
ifc_map_functions_acc.c |
Automatically-generated accessor function definitions for interacting
with the IFC modules format.
|
ifc_map_functions_dbg.c |
Automatically-generated debug functions for debugging objects of types
declared in
ifc_map.h. |
ifc_map_functions_mut.c |
Automatically-generated mutator function definitions for interacting
with the IFC modules format.
|
ifc_map_functions_val.c |
Automatically-generated validation function definitions for validating
objects of types declared in
ifc_map.h. |
ifc_modules.c |
Handling of operations common to reading and writing IFC module files.
|
ifc_modules_read.c |
Handling of reading IFC modules files.
|
ifc_modules_templ.c |
Template definitions and explicit instantiations shared by
ifc_modules.c and the IFC map functions files. |
ifc_modules_write.c |
Handling of writing IFC modules files.
|
il.c |
Low-level manipulation of intermediate language constructs.
|
il_alloc.c |
Allocation and initialization of intermediate language constructs.
|
il_display.c |
Display of the intermediate language in human-readable form.
|
il_read.c |
Reading the intermediate language from a file.
|
il_to_str.c |
Creating textual representations for certain IL entries
(e.g., constants, types).
|
il_walk.c |
Walking the in-memory intermediate language tree.
|
il_write.c |
Writing the intermediate language to a file.
|
inline.c |
Minimal inlining of function calls for IL lowering.
|
interpret.c |
IL interpreter for
constexpr functions. |
layout.c |
Laying out class objects.
|
lexical.c |
Source input and lexical scanning.
|
literals.c |
Conversion of literal constants to internal form.
|
lookup.c |
Name lookup.
|
lower_c99.c |
IL lowering from C99 to C89.
|
lower_eh.c |
IL lowering of exception handling constructs.
|
lower_il.c |
Intermediate language lowering (C++ constructs are translated to C
constructs).
|
lower_init.c |
IL lowering of initializations, new/delete, and
constructors/destructors.
|
lower_name.c |
Name mangling for IL lowering.
|
macro.c |
Definition of macros (
#define directive) and invocation
(expansion) of macros. |
mem_manage.c |
Memory management (dynamic allocation of storage).
|
modules.c |
Generic module file processing.
|
ms_attrib.c |
Processing of Microsoft attributes.
|
ms_metadata.cpp |
Processing of C++/CLI metadata from assemblies. Contains C++ source
code that can only be compiled on a Microsoft Windows platform.
|
overload.c |
Overload resolution for expression processing.
|
pch.c |
Precompiled header processing.
|
pragma.c |
Processing of
#pragma directives. |
preproc.c |
Preprocessing directives.
|
scope_stk.c |
Management of the scope stack.
|
src_seq.c |
Management of source sequence lists.
|
statements.c |
Statements scanning.
|
symbol_ref.c |
Processing for symbol references.
|
symbol_tbl.c |
Symbol table entry and look up, maintenance of name scoping.
|
sys_predef.c |
Definitions of system-specific predefined macros and assertions.
|
target.c |
Target configuration support.
|
templates.c |
Declaring and instantiating class and function templates.
|
trans_copy.c |
Copying of secondary translation unit IL to the primary IL.
|
trans_corresp.c |
Establishment of correspondences between externally-linked entities in
different translation units.
|
trans_unit.c |
Driver for translation-unit processing.
|
types.c |
Utility routines dealing with types.
|
unicode_name_fsm.c |
Automatically-generated data file for a finite state machine that
recognizes Unicode character names.
|
error_msg.txt |
Contains the definition of the error codes and error messages and is
used to generate
err_codes.h and err_data.h. See
mk_errinfo. |
error_tag.txt |
May be used to provide additional error “tags” or aliases that may be
used to reference a given error message in certain command line
options. This is used to generate
err_data.h. See
mk_errinfo. |
Note that in general the files are paired. A .h file containing
declarative information corresponds to a .c/.cpp file containing
executable code. The file name.h contains the declarations of the
data structures and externally-callable routines for the file name.c, and is, in effect, the definition of the external interface to that
source file.
There are also some .h files that contain configuration information and are
not paired with a .c file (e.g., targ_def.h).
In addition, there is a set of .h files that contain #includes of
other header files. Their purpose is to group header files – both to help
manage their often complicated interactions and to make better use of
compilation systems that perform header precompilation.
basic_hdrs.h |
Header files devoted to configuration.
|
fe_common.h |
Header files common to all
.c files. |
decl_hdrs.h |
Header files used for declaration processing.
|
expr_hdrs.h |
Header files used for expression processing.
|
lower_hdrs.h |
Header files used for IL lowering.
|
In rough numbers, the .h files together are about 320,000 lines, and the
.c files about 1,017,000 lines (including IL lowering, the IL display
utility, and the C- and C++-generating back ends). The utility programs are
another 15,000 lines, and the minimal runtime about 11,000 lines.
3.3. Coding Philosophy#
This section contains miscellaneous information on the coding style and philosophy used in writing the front end, which may be useful to those who must maintain it.
The configuration section of a previous chapter of this document contains some important information of this type.
Code should conform to the C++11 standard and not have any dialect-specific
constructs [1]. The environment dialect switches __ANSIC__,
__SYSV__, and __BSD__ should used to control inclusion of header files
that are not available in all versions of C++. basics.h (which should be
included first in every compilation) takes care of the most common header files
and the most common library routines (it will remap them via macros if
necessary).
The size argument on calls of functions like memcpy should be wrapped in a
call of the function-like macro size_t_arg if the size is anything other
than a sizeof operator.
Code that is host-dependent should be written only in host_envir.c and
cmd_line.c.
Target-dependent code should not be written outside of const_ints.c,
fixed_pt.c and float_pt.c (which are heavily target-dependent) and
literals.c, folding.c, and layout.c (which have some
target-dependent aspects). In particular, be careful about doing anything with
integers without going through the routines in const_ints.c.
3.3.1. External Names#
External routines should always be declared both where the body occurs and in
the corresponding .h file. The .h file should be included everywhere
the routine is referenced.
External variables should be declared in .h files with EXTERN as their
storage class. All .h files that contain external variables must be
included in fe_init.c; there, the macro EXTERN, which ordinarily is
mapped to extern, is mapped to a null string. This makes the same
declarations, when scanned in fe_init.c, external definitions instead of
external references. If an external variable should have an initial value, the
initialization should be included in the .h file, surrounded by #if
VAR_INITIALIZERS/#endif. In fe_init.c the initializer will be
included.
Of course, variables should be made static instead of external if it is
possible to localize them to a single source file.
All static variables (external, internal, and function-local) must be saved and
restored as necessary when precompiled headers are generated. The point at
which a precompiled header is written or read is outside of any declaration and
at the file scope (etc.); variables that never have interesting values at such
points need not be saved and restored. All others, however (and these cannot
be function-local), must be registered by calling
register_pch_saved_variables so that they will be written to the
precompiled header file and restored.
Likewise, variables that are translation-unit specific should be registered by
calling register_trans_unit_variable. See the Multiple Translation Units chapter.
3.3.2. Include Files#
Include files always include any other files on which they depend. They also
define a flag once they have been included, and will skip themselves if
included again. (For example, cmd_line.h defines CMD_LINE_H).
Therefore, one can include the files in any order and without regard for
whether or not they have been included by other include files. To improve
compilation speed, #includes can be surrounded by #ifndef
flag/#endif.
3.3.3. Debug Code#
Debug code is surrounded by #if DEBUG/#endif. In such code, one should
test the value of debug_level to control output of debug information.
debug_level is zero ordinarily, to indicate that no debug output should be
produced. Debug code should be set to come out at an appropriate level: 1
is for output that comes out once per source line, 3 for output that comes
out once per token, and 5 for output one would hardly ever want to see.
2 and 4 lie in between those levels in the expected way. Choosing an
appropriate debug level is more of an art than a science. If in doubt, make a
guess at the best level, run the front end with that level of debug output, and
see if the new messages seem out of place in the overall debug output at that
level.
One can also use db_flag_is_set to determine whether specific output should
be generated. For example, calls of db_flag_is_set("vtbl") are used in IL
lowering to control debug output related to virtual function tables. By
specifying -d-vtbl on the command line, one can get that particular debug
output and nothing else. A particular “flag” is considered set if the string
of the db_flag_is_set call is specified in a command-line debug option, so
a new “flag” is added by adding a call of db_flag_is_set with some new
string; there is no central registry of debug “flags.”
A similar facility for tracing names is provided by the --db_name=xxxx
command-line option and the db_has_traced_name macro. This can be used to
control extra output about processing of entities with specific names.
db_trace provides a combination of the features of db_flag_is_set and
db_has_traced_name.
Debugging options that are specified with the -d command-line option can
also be specified using the db_opt pragma. The pragma can be used to
specify a certain debug level or to define a certain debug flag for a portion
of a source program that is being compiled. For example:
template <class T> void f() {
#pragma db_opt -nondep_call
g();
#pragma db_opt #nondep_call
}
When a flag name is preceded by “#”, the flag is cleared if previously set.
Debugging options that are specified with the --db_name command-line option
can also be specified using the db_name pragma.
The variable db_active is TRUE if debug_level is non-zero or if there
is the potential for it becoming non-zero (i.e., if there is a list of debug
requests from the command line). This can be used to control processing that
may be needed when debug output is generated, but should not be done when no
debugging is requested.
Diagnostic messages should be written to f_error. This is normally set to
stderr, but in some modes is set to stdout, and can be redirected to a
file by use of a command-line option. Debug output should be written to
f_debug. This is set to f_error, but is a separate file variable to
allow it to be changed if desired.
Entries and exits of routines should have calls to db_enter and
db_exit. These provide flow tracing. The db_enter call looks like
db_enter(level, "routine-name");
The db_exit call has no arguments. These calls will cause an entry
message and an exit message when debug_level is equal to or greater
than the level specified. Note that the calls need not be surrounded by
#if DEBUG/#endif; db_enter and db_exit are actually macros,
and when DEBUG is 0, they are defined as null strings. When
DEBUG is non-zero, they expand to calls to debug_enter and
debug_exit. Those calls are surrounded by ifs testing
db_active so that the routine calls are only done when debug output is
actually being generated (or may be generated, when a list of
routine-related debug requests is specified on the command line).
It should be clear that for db_exit to work correctly one should have a
single exit point for each routine, and not have return statements in the
middle of routines.
Some routines do not have db_enter and db_exit calls for speed
reasons – even the overhead of two ifs would be visible. An example is
skip_white_space in lexical.c.
3.3.4. Checking Code#
Checking code performs dynamic consistency checking, e.g., verifying the entry
conditions for routines and the consistency of tables. Checking code is always
surrounded by #if CHECKING/#endif. When a problem is found, this code
calls
internal_error("routine-name:problem found");
which displays the argument string and causes abnormal termination of the front end.
As a concrete example of checking code: all switch statements should
have a default clause; if the default clause does not correspond to any
legal possibility, it should be present as checking code, and should call
internal_error to indicate that a bad value appeared as the switch
selector. (See CHECK_SWITCH_DEFAULT_UNEXPECTED in basics.h and
default_is_unexpected in checking.h for macros that assist in
applying this rule.)
Several routines are provided as convenient ways to do consistency
checking. Calls of these routines need not be surrounded by #if
CHECKING/#endif; they are defined as macros that expand to nothing
when CHECKING is disabled. check_assertion can be used to check an
expression to make sure it is true. expect_error checks that
total_errors is nonzero at the point of invocation or that it becomes
nonzero before the end of the compilation. This is useful to ensure that
situations that are expected to lead to error diagnostics will indeed do so
(only the first invocation will be reported on failure).
check_assertion_or_expect_error combines the previous two routines: it
has no effect if either the given condition is TRUE, or if total_errors
is nonzero by the end of the compilation. unexpected_condition is used
when a piece of code is not expected to be reached (for example, the
default clause of a switch). Three versions of each of these routines are
provided: the version described above, a version with a _str suffix
that accepts a single string argument, and a version with a _str2
suffix that accepts two strings. The string(s) are printed on error
termination to provide additional information about the cause of the error.
The consistency checks that are enabled by the CHECKING flag do not
significantly affect the performance of the front end. There are other
consistency checks whose execution does affect the performance of the front
end. These extra consistency checks are enabled by the EXPENSIVE_CHECKING
flag. These checks are not intended to be enabled when optimum performance is
required (e.g., in production versions of a product).
3.3.5. Conventions for Declarations#
The type a_boolean (defined by default as int) should be used for
true/false variables, along with the constants TRUE and FALSE.
These are defined in basics.h and thus are available everywhere. In
structs, a_byte_boolean is used for boolean fields when there is
just one of them. When there are several, bit fields of length one are
used, to save space. Obviously, these rules will not work on all host
machines, but they will work well enough on most.
Fields in structs are ordered to some extent to reduce the space wasted
in gaps for alignment. It’s assumed that pointers and large integers have
similar size and alignment requirements, and thus a sequence of them will leave
no gaps. When small fields appear, however, they are grouped together to avoid
the gaps one would expect from a series of short/long/short/long fields.
Bit fields are declared with the type a_bit_field. It is usually
unsigned int, but when compiling with a compiler that uses the Microsoft
bit-field layout algorithm it can be set to unsigned char to get better
packing of the bit fields.
The type a_byte should be used for small integer fields (i.e., fields
whose values lie in the range 0-127), rather than char or unsigned
char. enum types are often given a base type of a_byte so that
their size is explicitly controlled to save space in structs.
For each table type (whether or not it is part of the intermediate
language) there is an associated allocation routine and an associated
routine that will set the fields of the table to default values. (In some
cases, these two functions are provided by one routine.) When a table has a
variant section and an associated tag field, there is also a routine to set
the tag field and set the associated variant fields to default values. A
tag field should never be assigned directly; it should only be set by this
routine. Consistent use of the allocation, clear, and set-kind routines
ensures that no uninitialized fields will cause sporadic problems, and that
new fields can be easily added. In addition, it can be useful to add the
following snippet to defines.h :
#if UNION_AS_STRUCT
#define union struct
#endif
Setting UNION_AS_STRUCT in a testing build allocates each variant field
in its own storage. This often allows detection of cases where code
accesses an incorrect variant for a given IL entry.
The names of macros are written in upper case in the front end source when they are object-like (e.g., manifest constants), but in lower case when they are function-like. The names of types and enumerators are written in lower case.
The names of typedefed, class, and templated types always begin with
a_ or an_ (e.g., a_type, an_expr_node). Pointer type names
end in _ptr. The names of constants in enumerations always begin with
a prefix that hints at the enumeration type (e.g., enk_ for
an_expr_node_kind).
3.3.6. Storage Allocation#
Allocation of space is never done by calling malloc or the like directly.
For all table entries, as mentioned earlier, the associated allocation routine
should be called. It will allocate and clear the space, and will also make
sure that the space is allocated in the right area (e.g., intermediate language
constructs must be allocated in the right memory region). For any other
allocation cases, one of the routines in mem_manage.c should be called
instead. alloc_fe has an interface like malloc, and allocates storage
that need not survive the execution of the front end. alloc_general is
similar, and allocates storage that will survive into the back end if the back
end is called as part of the same program (this should not be used ordinarily).
One case to do carefully: all names that become part of the intermediate
language (file names, directory names, identifier names) should be allocated in
the file-scope intermediate language region, by calling alloc_il.
Space is never freed by calling free or related routines. Where tables are
reused, the front end maintains individual lists of available entries for each
type of table, and takes entries off those lists before allocating new ones.
Space can be released on a memory region basis, for example to free up all the
space used for a given function once the function has been processed. When the
front end is configured in such a way that it can be re-initialized (by setting
MAKE_FRONT_END_CALLABLE) the front end keeps a record of all memory
allocations and automatically frees the memory when the front end terminates.
In general, every reasonable attempt is made to not allocate tables unless they are actually needed. The intermediate language, for example, will contain some wasted space, but extremely little.
In order to avoid any fixed limits in the front end, the primary non-IL data
structures are designed so they can be dynamically allocated and expanded
(curr_source_line, scope_stack, etc.). An initial allocation is done
that should cover all expected uses; if more space is needed, the table is
reallocated with a larger size. That means that there is a danger inherent in
keeping around pointers into those data structures, because if the table moves,
all pointers to the table become obsolete. In the input line/macro area, such
pointers are necessary, and there is a pointer registration scheme that is used
to make sure that all pointers are updated if a reallocation occurs. In other
areas, pointers can be avoided relatively easily. A simple rule to remember is
that one should never make a local pointer into one of the primary stacks
(scope_stack, struct_stmt_stack); one should instead use an index value
into the stack, since an index is not invalidated by a reallocation of a table.
When memory is allocated that may later be resized, it must initially be
allocated using either alloc_resizable_buffer or realloc_buffer. It
may then be resized later by calling realloc_buffer.
One can get a summary of space use by having debug_level non-zero in
fe_wrapup. This can be requested by specifying -dfe_wrapup=1 on the
command line. A summary of symbol table and constant sharing efficiency is
also included.
3.3.7. Typerefs#
The description of a type in the intermediate language contains a type entry
kind tk_typeref, which is used to define typedefs, and also to add type
qualifiers to a type, e.g., const in const int. Such an entry is in
some ways not a full-fledged type. For example, it doesn’t indicate its size;
one must go to the underlying type to get that. The macro skip_typerefs
and the function f_skip_typerefs can be used to strip off typerefs and get
to the underlying type, which can then be processed in the normal way. A very
common cause of bugs in code our customers add is failure to put in
skip_typerefs calls where necessary. (For that matter, it’s a common cause
of bugs in our code, though we try hard to find all those bugs before we ship
the code.) The whole skip_typerefs issue is error-prone and requires
discipline, but you do get used to it, and after many years of working with it
we remain convinced that it’s the right way to handle typedefs and type
qualifiers. Just be careful.
3.3.8. Source Positions#
Source positions are represented as a structure of type a_source_position
(see basics.h). It contains a sequence number and a column number. The
sequence number is a unique number assigned in ascending order to each source
line as it is read. Unlike a file name and line number, the sequence number is
unambiguous. It’s also more compact, and can be converted to a file name and
line number for external use (this is done in error messages, for example).
The column number is a 1-origined character position in the original source line. The character position is the logical character number in the line, not simply a byte offset. This is significant if the source file contains multibyte characters or is in a format such as UTF-16. Tab characters are counted as a single column. The position of text within a macro expansion is considered to be the first character of the macro name in the macro call in the original source line that directly or indirectly generated the text.
If the front end is configured with FULLY_RESOLVED_MACRO_POSITIONS set to
TRUE, each source position contains an additional sequence/column number pair.
For text within a macro expansion, this additional position information gives
the location from which the associated text was copied into the macro
expansion, i.e., a macro definition or macro argument.
If MACRO_INVOCATION_TREE_IN_IL is configured to TRUE, each source position
referring to text within a macro expansion contains the index of the macro
invocation record that identifies the macro in whose expansion the text at that
position occurs. For text appearing directly in the source, this index will be
NO_PARENT_MACRO_INVOCATION.
A sequence number of zero is used to indicate an unknown position or a
“position” in initialization, in a predefined macro, or in the command-line.
The column number indicates which. See basics.h.
Importing assembly metadata in C++/CLI is a process that creates a string
containing C++ code that expresses that metadata and then parses the contents
of that string. Each import is associated with a particular sequence number,
and that number is assigned to the position of each token extracted from that
string (with a column number zero). For example, when compiling a C++/CLI
source file, sequence number one will usually correspond to mscorlib.dll.
3.3.9. Language Dialect#
The global variable C_dialect (in cmd_line.h) indicates the C
dialect to be accepted: the choices are C_dialect_cplusplus,
C_dialect_ANSI, and C_dialect_pcc. The macro C_mode() is
useful as a way of testing for C versus C++. cppNN_mode
(where NN is the year of the relevant C++ Standard, e.g., cpp03_mode,
cpp17_mode, etc.) is TRUE when the features of the corresponding
revision of the language should be accepted, and similarly cNN_mode for versions of the C language.
The global variable strict_ansi_mode is TRUE if use of features that
are not part of the selected C or C++ Standard should cause diagnostics.
strict_ansi_error_severity indicates the level of diagnostic (warning,
discretionary error, or error) that should be produced for such cases.
microsoft_mode indicates whether a Microsoft dialect (C, C++, C++/CLI,
or C++/CX) is in effect. In configurations that do not allow Microsoft
extensions (i.e., when MICROSOFT_EXTENSIONS_ALLOWED is FALSE), it is a
macro equivalent to FALSE. This allows many Microsoft-mode-specific bits
of code to be written without surrounding them with conditional compilation
directives testing MICROSOFT_EXTENSIONS_ALLOWED. cppcli_enabled
and cppcx_enabled similarly indicate that the Microsoft dialect is
C++/CLI or C++/CX, respectively (if TRUE, microsoft_mode is TRUE too,
and C_mode() is FALSE). Some C++/CLI- and C++/CX-specific code is
delimited by conditional compilation directives testing
MICROSOFT_EXTENSIONS_ALLOWED because it uses variables, data members,
or functions that are available only when that macro is TRUE. (The macro
CPPCLI_ENABLING_POSSIBLE is tested in a few places for code that deals
with the importing of assembly metadata, but it shouldn’t be used to
conditionalize other C++/CLI-specific code. In particular, the IL
extensions representing C++/CLI and C++/CX constructs are only conditioned
on MICROSOFT_EXTENSIONS_ALLOWED.)
gpp_mode and gcc_mode indicate that, respectively, GNU C++ or GNU C
emulation is in effect. In configurations where GNU_EXTENSIONS_ALLOWED is
FALSE, these are macros equivalent to FALSE.
clang_mode indicates that clang emulation is in effect. Because the
initial emulation of the clang C++ compiler in the front end was very close
to g++, clang_mode also implies that gpp_mode or gcc_mode will
be set. When it is necessary to distinguish between the GNU and clang
emulations, the macros gnu_version_is, gpp_version_is,
gcc_version_is, clang_version_is, clangcpp_version_is, and
clangc_version_is (see lang_feat.h) allow finer-grained
discrimination (including the compiler version to be emulated). (There are
also corresponding Microsoft-specific macros, ms_version_is,
mscpp_version_is, and msc_version_is.)
cfront_2_1_mode is TRUE if the cfront 2.1 dialect of C++ is to be compiled.
cfront_3_0_mode indicates the cfront 3.0 dialect. The macro
any_cfront_mode() tests for either of those being set to TRUE.
(There are other variables/macros that control dialect variations. The ones mentioned here are among the more frequently used ones.)
3.3.10. Expression Processing#
See Rescanning Expressions for some additional coding rules that apply only within the expression-processing routines.
3.4. Recursive Descent Parsing#
“Recursive descent” is an ad-hoc syntax parsing technique. It works well for languages (like C) where very little lookahead is required to resolve local ambiguities, and where the language in which the compiler is written is recursive (as is true for the front end).
The idea of recursive descent provides a framework and a discipline for parsing a language, but it is implemented entirely by hand-crafted code. There is nothing automatic about the parsing. Thus, there is more possibility for error in developing a recursive-descent parser than there is in developing a table-driven parser. However, if one writes carefully and tests thoroughly, the end result is better with a recursive descent parser: faster parsing, better error recovery, and a clearer compiler (because all of the logic related to a particular construct appears in one place).
To implement recursive descent, one writes routines for each non-terminal construct in the language syntax. For example, one writes a routine to scan an expression, and one to scan a statement, and one to scan a declaration. The internal structure of these routines looks a great deal like the syntax of the constructs they implement. For example, if a declaration is defined as a list of declaration specifiers followed by a declarator followed by a semicolon, the routine that scans a declaration would call the routine to scan a list of declaration specifiers, then the routine to scan a declarator, and then check for and discard a semicolon. Where the language has recursive constructs, the routines involved will call each other recursively.
The terminal tokens that are being parsed come from the lexical routines,
specifically from get_token. At any given moment, there is a current
token (curr_token). It is the next token of input, the one that
hasn’t yet been taken. If the token is an identifier or a literal
constant, there is additional information associated with it. See
lexical.h.
As the recursive descent routines parse the input tokens, they build up information on what they have parsed, in the intermediate language trees and in the symbol table. They also check for and recover from errors.
The scanning of expressions is one area where recursive descent is inefficient
for languages (like C) where the expression syntax has many levels. A naive
recursive descent implementation would provide a routine for each level of the
expression syntax, and thus scanning a simple expression would involve perhaps
a dozen subroutine calls. For this reason, the front end uses a modified
recursive descent technique to scan expressions. Internal calls to the
scan_expr routine (in expr.c) specify a precedence level, and the scan
routine can decide whether to call itself recursively or not based on the
relative precedences of adjacent operators. With this modification, the number
of calls is reduced to very nearly one per operator. See the description of
expr.c later in this document for more detail.
To reduce the possibilities for errors in writing the parsing routines, one
should not check for the start of constructs by checking directly for all the
possible first tokens in that construct. Such a list of tokens is inherently
error-prone. For constructs of this kind, there are functions that will
indicate whether or not the current token looks like the start of the construct
(see, for example, is_decl_start in decls.c). Such routines might
themselves call other is_xxx_start routines if the constructs can
themselves begin with other complex constructs.
Several utility routines in lexical.c are useful for recursive descent
parsing. required_token can be used to check for and discard a required
token, and to issue a specific error message if the token is not present.
loop_token is used to check (at the bottom of the loop implementing an
iterative syntactic construct) for the occurrence of a particular token. The
token is taken if found, but no error is given if the token is not found.
next_token is useful for doing a limited kind of lookahead: it will return
the kind of the token following the current token, without actually getting
that token (actually, it gets it and then puts it back). In many cases, that
one token is enough to allow a decision to be made at a fork in the syntax.
See, for example, the code that recognizes statement labels in statement in
statements.c.
Consult any of the popular books on compiler theory for a more detailed overview of recursive descent.
3.5. Error Recovery#
When the front end discovers an error, it has several goals:
- To describe the error clearly and indicate its position (including the position within the source line);
- to recover from the error without generating extra errors;
- to preserve as much information as possible about the valid parts of the construct containing the error, so that detection of other errors is not needlessly compromised.
(“Error” here of course refers to remarks and warnings as well.)
Describing an error clearly requires that one not only detect the error, but
detect it at an appropriate level, so that the front end can guess better at
what was intended. Sometimes this requires clever code at decision points in
recursive descent parsing, so that on invalid constructs one goes to the more
likely of two subcases. Within the { } for a struct, for example, we
have a series of member declarations separated by semicolons. If a semicolon
is missing, one should not fall out of the loop and leave the scanning of the
struct. In other situations, clearly invalid cases should be flagged at
the higher level, and none of the subcases should be entered. For example, at
the beginning of a statement or a declaration, if the first token is not a
valid start for that construct, an error message (“expected a statement” or
“expected a declaration”) is more informative than descending into any of the
subcases and letting them issue errors.
3.5.1. Error Positions#
The position of an error is often just as important as the error text. The
front end tries to give useful error positions. This is aided by the fact that
column positions are maintained for each token. A global variable
error_position indicates the position of the beginning of the last language
construct scanned. It is set by get_token on scanning each token, and it
is also set by each medium-sized language construct (e.g., an expression, a
declarator). Most recursive descent routines save the position of the first
token of the construct they scan, and then set error_position to that
position on return. error_position is the default position for errors when
no explicit position is specified, and this protocol for setting it works well
in ensuring that the error positions are right. Some cases take extra work.
For example, in the expression routines, the position of the left operand and
right operand, and the position of the operator itself, are known for each
operation. If an error is due to a problem with one of the operands, the error
position used is the position of that operand. If the problem is with the
combination of those operands (as in adding a pointer and a float), the error
position used is the position of the operator.
The macro set_err_pos_to_curr_token is used when one is going forward in
scanning. It sets the error position to the position of the current token.
One would use it, for example, when looking for the = that indicates an
initializer in a declaration. The previous construct (a declarator) has been
scanned and checked, and its position is in error_position. Now, however,
we are going forward, and the = that may have ended the expression scan
should once again be the current construct. In practice, there are few cases
like this; they only come up for optional tokens in the middle of constructs.
required_token handles the change of error position in the more usual case
of a required token.
When an error refers back to something earlier in the source, it often makes sense to give the position of the other location rather than the current location. Variables declared but never used, for example (a warning), are detected at the closing brace of functions. However, that position is much less informative than the position of the declaration of the variable in question, so that is used. Likewise, when a comment is unclosed at end of file, it is more useful to indicate the position of the start of the comment instead of the location of the end of file, where the error was detected.
3.5.2. Error Codes#
Error messages are mapped to values of the enumeration an_error_code (see
err_codes.h, along with err_data.h for the associated error message
text). For example, ec_exp_line_number represents the message “expected a
line number”. ec_no_error is first in the enumeration, and therefore has
the value zero. Thanks to it, one can return error codes of type
an_error_code, and test them in if statements (non-zero means an
error).
Errors with codes beginning ec_exp_ (and texts beginning “expected a …”)
are uniformly syntax errors – some expected token or construct was not found.
The severity of an error is determined by the error routine called (e.g.,
error instead of warning) and not by the error code. In fact, there
are cases where the same error code is used as both a warning and an error.
The severity of certain diagnostic messages may be overridden using a command
line option. A diagnostic that is issued with a severity that is less than an
error (i.e., remark, warning, or discretionary error) may have its severity
changed, or may be suppressed completely.
Some error message texts have fill-ins, as in
could not open source file "name"
For these, variants of the error routines should be called. For the case
above, for example, str_error would be called. The fill-in string is
passed to the error routine as a null-terminated string, and replaces
%s where it appears in the error text string. See
Error Reporting for details on the kinds of fill-ins available.
3.5.3. Syntax Errors#
In terms of internal processing, one must distinguish syntax errors from
semantic errors. Syntax errors are discovered when the sequence of tokens
in the source program cannot be parsed according to the language syntax
(for example, i=j+;–there’s something missing between the + and
the ;). All other errors are semantic errors – things that can be
parsed but are still nonsensical (for example, use of an undefined
variable, or adding two structs). In error recovery, the important
distinction between the two is that for a syntax error one must decide how
to continue parsing, whereas for a semantic error everything is fine as far
as parsing and there’s no question about how to continue.
Syntax errors should always cause the calling of flush_tokens in
lexical.c, and semantic errors should never call it. The routine
syntax_error records the occurrence of an error, and then calls
flush_tokens, which is a convenient combination.
flush_tokens scans and throws again tokens until it finds one in the
stop tokens set, and then returns. Paired tokens are also considered: if
a ( is scanned while flushing tokens, for example, the corresponding
) will be sought before continuing the search for a stop token. There
are some other special cases; see the code in lexical.c.
The stop tokens set is an array indexed by token kind. Each element of the
array is non-zero if the corresponding token is currently in the stop tokens
set (and thus should halt a flush). As recursive descent routines scan
constructs, they add to the stop tokens set those terminal tokens they expect
to see in their constructs, and then remove those tokens once they are
encountered. The macros add_stop_token and remove_stop_token are
called to add and remove tokens from the set by incrementing and decrementing
the associated element of the array. Since syntactic constructs are nested
inside one another, having each recursive descent routine add the tokens it
recognizes produces a stop tokens set that will halt a flush at a token with
which one of the active routines feels it can continue.
At boundaries where the syntax changes dramatically, push_stop_token_stack
is called to establish a new set of stop token values. The old values are
restored later by calling pop_stop_token_stack. This is done, for example,
when entering the processing for preprocessing directives, whose syntax is
dramatically different than the context in which they are embedded.
Uses of add_stop_token and remove_stop_token must be very carefully
paired. A missing call of either would probably cause the associated token to
be handled incorrectly in error flushes for the rest of the compilation.
Skipping a remove_stop_token call is easy to do if one does a goto out
of an inner loop in a recursive descent routine (say, because there is an
error) and one skips over the point where the remove_stop_token is done.
The likelihood of such a problem is lessened by having debug_enter and
debug_exit verify (by means of a checksum) that the contents of
stop_token_array are the same on exit as they were on entry. This check is
only done when debugging is active, however, because it’s time-consuming.
Also, fe_wrapup verifies that all the counts have returned to zero by the
end of the compilation; that check is always done. In addition, when expensive
checking code is enabled, pop_stop_token_stack verifies that all of the
elements of the stop token set that is being discarded have been reset to zero.
If required_token does not find the token it expects, it adds that token to
the stop tokens set, and calls syntax_error. After flush_tokens
finishes, required_token checks again to see if the desired token has come
up, and if so, scans over it.
3.5.4. Error Entries#
For both syntax and semantic errors, another part of error recovery is producing something sensible, in spite of the error, in the data structure that represents the program. This task is eased by having special “error” entries for all of the major table types. There are, for example, error constants, error types, and error expression nodes. They are used to take the place of any part of the program tree that cannot be built because of an error. In general, the code tries to build as much of the complete tree as possible. An error in one part of an expression, for example, need not mean that the entire expression tree will be replaced by an error expression node. Throughout the front end, code must be capable of dealing with error entries. The basic philosophy is that an error entry is considered to be compatible with anything else. No additional errors should be issued on encountering an error entry, since it is assumed that some error was issued when the entry was created.
When syntax errors involve missing identifiers, special error symbols are
created. A different error symbol is created for each error case. All error
symbols point to the same special associated symbol header (see the
documentation of symbol_tbl.c).
When errors involve duplicate declarations of the same symbol, the second symbol is entered anyway. Both symbols are in the symbol table, but because of the symbol table structure, the later declaration will the one that is visible, and it therefore effectively replaces the earlier declaration.
In general, scanning routines do not need to return indications of success or failure, since they always return a tree representing the construct scanned. Some routines must deal with unusual cases where there is no possible representation for the thing scanned, and so they return an error parameter. The caller must check it.
3.5.5. Intelligent Error Diagnosis#
The front end attempts to diagnose as many errors and warnings as possible. If
one uses header files with function prototypes for all external functions and
variables, it should be possible when writing in ANSI C or in C++ to do away
with lint, or at least to relegate it to use in unusual cases. That is the
goal of the front end: diagnose not only errors but also questionable
constructs such as those flagged by lint. However, since lint is run
on request and the front end must be run on each compilation, one must take
care that the errors are not overly pedantic. The front end tries to do this
by being selective about errors. It’s a good idea, for example, to warn about
value-less return statements in functions that should return a value.
However, in old-style C, before the keyword void existed, it was standard
practice to omit the type specifier altogether on a function that returned
nothing. The C language says such a function has type int, but the C
programmer usually thinks of such a function as having type void. So, in
the front end, a warning is issued for a value-less return in a function with
an explicit type, but not for one in a function with an implicit int type.
In pcc mode, a warning is never issued.
A similar issue comes up in expressions. An expression like 1/0 looks like
an error (division by zero), but should not necessarily cause a compilation
error. If it occurs in a context where it is not evaluated (such as
0&&1/0), it never causes an error; if it occurs in a context where it must
be evaluated at compile time (as in the size of an array), it causes an error;
and if it occurs elsewhere (i.e., in executable code), it causes a warning and
is left as-is to cause a fault (or not) at execution time.
Likewise, when a constant seems to be part of a bit-oriented operation (it is a non-decimal constant, or it results from a bit-wise operator applied to constant operands), the normal warnings about implicit sign changes when the constant is changed from signed to unsigned or vice-versa are suppressed.