3. Overview of Internal Structure#

3.1. Function of the Front End#

The front end reads source files written in C or C++, diagnoses any errors in them, and produces an intermediate language tree that represents the source program.

The source files are one primary source file (named on the command line) and zero or more header files (specified by #include directives). If COMPILE_MULTIPLE_SOURCE_FILES or COMPILE_MULTIPLE_TRANSLATION_UNITS is TRUE, multiple source files may be specified on the command line.

The intermediate language tree is an in-memory representation of the source program, from which a back end can generate object code. This intermediate form is high-level (i.e., its constructs correspond closely to C or C++ constructs). Implicit transformations in the source language are made explicit in the intermediate language. Source-correspondence information is preserved so that symbolic debugging information can be generated.

When it translates C programs to intermediate language, the front end uses only the C subset of the intermediate language. When it translates C++ programs to intermediate language, it uses the full intermediate language. However, an optional IL lowering pass can be used translate the C++ intermediate language constructs into C intermediate language constructs (this lowering is not available for C++/CLI and C++/CX).

The front end does not do any optimizations. It does not transform the program in other than trivial ways.

The front end would usually be run from a driver, which invokes the front end and back end in order, and which validates and distributes the command-line options. In such a case, the front end would write the intermediate language information to a file, which would be read by the back end.

The front end incorporates the preprocessing function of the C and C++ languages, and does them on the fly as it parses the source program. The front end can also be run in a preprocessing-only mode, and will generate preprocessing output as a separate preprocessor would. Comments can be kept or removed in this process.

The front end can generate a file containing raw listing information to be used to assemble a compilation listing. The raw listing file contains raw source lines, information on transitions into and out of include files, and diagnostics generated by the front end. A file containing cross-reference information (information on each reference to an identifier in the source program) can also be generated in a different file.

3.2. Guide to the Source Files#

Note that most source files have a .c file suffix, even though they contain C++ code and must be compiled with a C++ compiler (with the C++11 dialect). That naming convention is for historical reasons and for ease in delivering patches against earlier versions of the front end. Symbolic links (e.g., with a .cpp suffix) can be added if required by a build environment.

The source files are:

attribute.h
Declarations related to attribute.c (scanning of attributes).
basics.h
Basic declarations for every compilation.
builtin_defs.h
Declarations related to GCC/clang/Microsoft builtin functions. This header file is generated automatically by a tool that retrieves the signatures of these builtin functions from various versions of GCC and clang compilers (the Microsoft builtin function signatures are manually input to this automated process).
builtin_kinds.h
Builtin function kinds (split from builtin_defs.h to improve compilation times). This header file is generated automatically by a tool that retrieves the signatures of these builtin functions from various versions of GCC and clang compilers (the Microsoft builtin function signatures are manually input to this automated process).
c_gen_be.h
Declarations related to c_gen_be.c (C-generating back end).
cfe_daemon_common.h
Declarations related to cfe_daemon.c exposed for utility programs.
checking.h
Fundamental declarations associated with assertion checking.
class_decl.h
Declarations related to class_decl.c (class declarations scanning).
cmd_line.h
Declarations related to cmd_line.c (command-line parsing).
const_ints.h
Declarations related to the manipulation of the internal representation of integer constants.
cp_gen_be.h
Declarations related to cp_gen_be.c (C++/C-generating back end).
debug.h
Declarations related to debug.c (debug output).
decl_inits.h
Declarations related to decl_inits.c (declaration initializers scanning).
decl_spec.h
Declarations related to decl_spec.c (declaration specifiers scanning).
declarator.h
Declarations related to declarator.c (declarator scanning).
decls.h
Declarations related to decls.c (declarations scanning).
def_arg.h
Declarations related to def_arg.c (default argument scanning).
defines.h
A place to put #defines to configure the front end; defines.h is included at the beginning of every source file.
direct_allocator.h
Declaration and definition of the Direct_allocator template.
disambig.h
Declarations related to disambig.c (C++ expression-declaration disambiguation).
error.h
Declarations related to error.c (error reporting).
err_codes.h
Definition of the enumeration of error codes. This file is generated from the error_msg.txt file. See mk_errinfo.
err_data.h
Definition of the enumeration of error codes. This file is generated from the error_msg.txt and error_tag.txt files. See mk_errinfo.
expr.h
Declarations related to expr.c (expression scanning).
exprutil.h
Declarations related to exprutil.c (expression scanning utilities). Also, the data structures used in expression scanning.
extasm.h
Declarations related to extasm.c (GNU C extended asm statement scanning).
fe_init.h
Declarations related to fe_init.c (initialization).
fe_wrapup.h
Declarations related to fe_wrapup.c (termination).
fixed_pt.h
Declarations related to fixed_pt.c (manipulation of fixed-point constants).
float_pt.h
Declarations related to float_pt.c (manipulation of floating-point constants).
float_type.h
Definitions of floating-point conversion routines used when USE_HOST_FP_CONVERSION_ROUTINES is FALSE. This header file is included multiple times (once for each distinct floating-point type) by floating.c.
floating.h
Declarations related to floating.c (software-based conversion of floating-point values between binary and decimal).
folding.h
Declarations related to folding.c (folding of constant operations).
func_def.h
Declarations related to func_def.c (processing for function definitions).
header_util.h
General utility components (mostly templates) intended for use in both headers and compilation units.
host_envir.h
Declarations related to host_envir.c (definition of the host environment, including operating system dependencies). Also, many host configuration switches.
host_util.h
Definitions of host environment functions that are used by both the front end and utility programs such as the prelinker. When the front end is compiled, these routines are included as part of host_envir.c.
ifc_map.h
Automatically-generated type declarations and definitions for interacting with the IFC modules format.
ifc_map_functions.h
Automatically-generated function declarations for interacting with the IFC modules format.
ifc_modules_inst.h
Automatically-generated macro invocations to facilitate instantiations in ifc_modules_templ.c.
ifc_modules_internal.h
Declarations related only to the internal implementation files for IFC modules file support.
ifc_modules_spec.h
Automatically-generated macro invocations to facilitate explicit specializations in ifc_modules_templ.c.
ifc_modules.h
Declarations related to IFC modules files.
il.h
Declarations related to il.c (intermediate language).
il_alloc.h
Declarations related to il_alloc.c (allocation of intermediate language entities).
il_def.h
The definition of the intermediate language.
il_display.h
Declarations related to il_display.c (display of intermediate language in human-readable form).
il_file.h
Declarations related to the structure of the intermediate language file.
il_read.h
Declarations related to il_read.c (reading the intermediate language file).
il_to_str.h
Declarations related to il_to_str.c (creating external string-form representations for certain IL entries).
il_walk.h
Declarations related to il_walk.c (walking the intermediate language tree).
il_write.h
Declarations related to il_write.c (writing the intermediate language file).
inline.h
Declarations related to inline.c (minimal inlining for IL lowering).
interpret.h
Declarations related to interpret.c (IL interpreter for constexpr functions).
lang_feat.h
Declarations of configuration switches that affect the source language accepted.
layout.h
Declarations related to layout.c (class object layout).
lexical.h
Declarations related to lexical.c (source input and lexical scanning). Also, the definition of the data structures related to the current input line.
lint.h
A place to put directives to silence the lint program.
literals.h
Declarations related to literals.c (conversion of literal constants).
lookup.h
Declarations related to lookup.c (name lookup).
lower_c99.h
Declarations related to lower_c99.c (C99 IL lowering).
lower_eh.h
Declarations related to lower_eh.c (IL lowering of exceptions).
lower_il.h
Declarations related to lower_il.c (IL lowering).
lower_init.h
Declarations related to lower_init.c (IL lowering of initializations).
lower_name.h
Declarations related to lower_name.c (name mangling for IL lowering).
macro.h
Declarations related to macro.c (macro definition and invocation).
mem_manage.h
Declarations related to mem_manage.c (memory management).
mem_tables.h
Declarations related to the memory management data structure.
modules.h
Declarations related to the processing of module files.
ms_attrib.h
Declarations related to Microsoft attribute processing.
ms_metadata.h
Declarations related to ms_metadata.cpp (reading C++/CLI metadata from assemblies).
overload.h
Declarations related to overload.c (overload resolution for expression processing).
pch.h
Declarations related to pch.c (precompiled header processing).
pragma.h
Declarations related to pragma.c (processing of #pragma directives).
preproc.h
Declarations related to preproc.c (preprocessing directives).
scope_stk.h
Declarations related to scope_stk.c (scope stack management).
src_seq.h
Declarations related to src_seq.c (source sequence list management).
statements.h
Declarations related to statements.c (statement scanning). Also, the definition of the statement stack.
symbol_ref.h
Declarations related to symbol_ref.c (symbol references).
symbol_tbl.h
Declarations related to symbol_tbl.c (symbol table). Also, the definitions of the symbol table and the scope stack.
sys_predef.h
Declarations related to sys_predef.c (system-specific predefined macros). Also declarations related to builtins (and the place to add user-defined builtins).
targ_def.h
Declarations that define default target machine characteristics, and target configuration switches.
target.h
Declarations of target configuration variables, initialized to default values defined in targ_def.h.
target_cfg.h
Definitions of target-specific routines to set and dump target configurations (included once for each target configuration).
target_map.h
Definition of various routines that map target-specific configuration macros to global variables in some fashion (included multiple times).
templates.h
Declarations related to templates.c (class and function templates).
trans_copy.h
Declarations related to trans_copy.c (copying of secondary translation unit IL to primary IL).
trans_corresp.h
Declarations related to trans_corresp.c (establishment of correspondences between translation units).
trans_unit.h
Declarations related to trans_unit.c (driver for translation-unit processing).
types.h
Declarations related to types.c (utility routines dealing with types).
util.h
General utility components. This contains a set of (mostly templated) utilities that are used in place of those that would normally be found in a C++ standard library. The declarations are conditionally enclosed in the edg namespace (under control of the USE_EDG_NAMESPACE configuration macro).
version.h
Definition of the current version number of the front end.
walk_entry.h
Routines used by il_walk.c to walk through IL entries.
attribute.c
Scanning of attributes.
c_gen_be.c
The C-generating back end (for use in place of a “real” back end; generates C code which can be compiled by a native target compiler).
cp_gen_be.c
The C++/C-generating back end (for use in source-to-source transformation).
cfe.c
The main program for the front end.
cfe_daemon.c
The main program for the front end when built as a daemon..
class_decl.c
Class declarations scanning.
cmd_line.c
Command-line processing.
const_ints.c
Manipulation of the internal representation of integer constants.
debug.c
Debug control routines.
decl_inits.c
Declaration initializers scanning.
decl_spec.c
Declaration specifiers scanning.
declarator.c
Declarator scanning.
decls.c
Declarations scanning.
def_arg.c
Default argument caching and rescanning.
disambig.c
C++ expression-declaration disambiguation.
error.c
Error reporting.
expr.c
Expression scanning.
exprutil.c
Utilities for expression scanning.
extasm.c
Scanning of GNU C extended asm statement.
fe_init.c
Overall initialization.
fe_wrapup.c
Overall termination.
fixed_pt.c
Manipulation of fixed-point constants: conversion to and from internal form and conversions and operations on the internal form.
float_pt.c
Manipulation of floating-point constants: conversion to and from internal form and conversions and operations on the internal form.
floating.c
Conversion of floating-point values between binary and decimal formats. Used only when USE_HOST_FP_CONVERSION_ROUTINES is FALSE.
folding.c
Folding of constant operations.
func_def.c
Processing for function definitions.
host_envir.c
Handling of host-environment dependencies (e.g., file name rules and opening of files).
ifc_map_functions.c
Automatically-generated function definitions for interacting with the IFC modules format (that do not fit into one of the more specific IFC map functions files).
ifc_map_functions_acc.c
Automatically-generated accessor function definitions for interacting with the IFC modules format.
ifc_map_functions_dbg.c
Automatically-generated debug functions for debugging objects of types declared in ifc_map.h.
ifc_map_functions_mut.c
Automatically-generated mutator function definitions for interacting with the IFC modules format.
ifc_map_functions_val.c
Automatically-generated validation function definitions for validating objects of types declared in ifc_map.h.
ifc_modules.c
Handling of operations common to reading and writing IFC module files.
ifc_modules_read.c
Handling of reading IFC modules files.
ifc_modules_templ.c
Template definitions and explicit instantiations shared by ifc_modules.c and the IFC map functions files.
ifc_modules_write.c
Handling of writing IFC modules files.
il.c
Low-level manipulation of intermediate language constructs.
il_alloc.c
Allocation and initialization of intermediate language constructs.
il_display.c
Display of the intermediate language in human-readable form.
il_read.c
Reading the intermediate language from a file.
il_to_str.c
Creating textual representations for certain IL entries (e.g., constants, types).
il_walk.c
Walking the in-memory intermediate language tree.
il_write.c
Writing the intermediate language to a file.
inline.c
Minimal inlining of function calls for IL lowering.
interpret.c
IL interpreter for constexpr functions.
layout.c
Laying out class objects.
lexical.c
Source input and lexical scanning.
literals.c
Conversion of literal constants to internal form.
lookup.c
Name lookup.
lower_c99.c
IL lowering from C99 to C89.
lower_eh.c
IL lowering of exception handling constructs.
lower_il.c
Intermediate language lowering (C++ constructs are translated to C constructs).
lower_init.c
IL lowering of initializations, new/delete, and constructors/destructors.
lower_name.c
Name mangling for IL lowering.
macro.c
Definition of macros (#define directive) and invocation (expansion) of macros.
mem_manage.c
Memory management (dynamic allocation of storage).
modules.c
Generic module file processing.
ms_attrib.c
Processing of Microsoft attributes.
ms_metadata.cpp
Processing of C++/CLI metadata from assemblies. Contains C++ source code that can only be compiled on a Microsoft Windows platform.
overload.c
Overload resolution for expression processing.
pch.c
Precompiled header processing.
pragma.c
Processing of #pragma directives.
preproc.c
Preprocessing directives.
scope_stk.c
Management of the scope stack.
src_seq.c
Management of source sequence lists.
statements.c
Statements scanning.
symbol_ref.c
Processing for symbol references.
symbol_tbl.c
Symbol table entry and look up, maintenance of name scoping.
sys_predef.c
Definitions of system-specific predefined macros and assertions.
target.c
Target configuration support.
templates.c
Declaring and instantiating class and function templates.
trans_copy.c
Copying of secondary translation unit IL to the primary IL.
trans_corresp.c
Establishment of correspondences between externally-linked entities in different translation units.
trans_unit.c
Driver for translation-unit processing.
types.c
Utility routines dealing with types.
unicode_name_fsm.c
Automatically-generated data file for a finite state machine that recognizes Unicode character names.
error_msg.txt
Contains the definition of the error codes and error messages and is used to generate err_codes.h and err_data.h. See mk_errinfo.
error_tag.txt
May be used to provide additional error “tags” or aliases that may be used to reference a given error message in certain command line options. This is used to generate err_data.h. See mk_errinfo.

Note that in general the files are paired. A .h file containing declarative information corresponds to a .c/.cpp file containing executable code. The file name.h contains the declarations of the data structures and externally-callable routines for the file name.c, and is, in effect, the definition of the external interface to that source file.

There are also some .h files that contain configuration information and are not paired with a .c file (e.g., targ_def.h).

In addition, there is a set of .h files that contain #includes of other header files. Their purpose is to group header files – both to help manage their often complicated interactions and to make better use of compilation systems that perform header precompilation.

basic_hdrs.h
Header files devoted to configuration.
fe_common.h
Header files common to all .c files.
decl_hdrs.h
Header files used for declaration processing.
expr_hdrs.h
Header files used for expression processing.
lower_hdrs.h
Header files used for IL lowering.

In rough numbers, the .h files together are about 320,000 lines, and the .c files about 1,017,000 lines (including IL lowering, the IL display utility, and the C- and C++-generating back ends). The utility programs are another 15,000 lines, and the minimal runtime about 11,000 lines.

3.3. Coding Philosophy#

This section contains miscellaneous information on the coding style and philosophy used in writing the front end, which may be useful to those who must maintain it.

The configuration section of a previous chapter of this document contains some important information of this type.

Code should conform to the C++11 standard and not have any dialect-specific constructs [1]. The environment dialect switches __ANSIC__, __SYSV__, and __BSD__ should used to control inclusion of header files that are not available in all versions of C++. basics.h (which should be included first in every compilation) takes care of the most common header files and the most common library routines (it will remap them via macros if necessary).

The size argument on calls of functions like memcpy should be wrapped in a call of the function-like macro size_t_arg if the size is anything other than a sizeof operator.

Code that is host-dependent should be written only in host_envir.c and cmd_line.c.

Target-dependent code should not be written outside of const_ints.c, fixed_pt.c and float_pt.c (which are heavily target-dependent) and literals.c, folding.c, and layout.c (which have some target-dependent aspects). In particular, be careful about doing anything with integers without going through the routines in const_ints.c.

3.3.1. External Names#

External routines should always be declared both where the body occurs and in the corresponding .h file. The .h file should be included everywhere the routine is referenced.

External variables should be declared in .h files with EXTERN as their storage class. All .h files that contain external variables must be included in fe_init.c; there, the macro EXTERN, which ordinarily is mapped to extern, is mapped to a null string. This makes the same declarations, when scanned in fe_init.c, external definitions instead of external references. If an external variable should have an initial value, the initialization should be included in the .h file, surrounded by #if VAR_INITIALIZERS/#endif. In fe_init.c the initializer will be included.

Of course, variables should be made static instead of external if it is possible to localize them to a single source file.

All static variables (external, internal, and function-local) must be saved and restored as necessary when precompiled headers are generated. The point at which a precompiled header is written or read is outside of any declaration and at the file scope (etc.); variables that never have interesting values at such points need not be saved and restored. All others, however (and these cannot be function-local), must be registered by calling register_pch_saved_variables so that they will be written to the precompiled header file and restored.

Likewise, variables that are translation-unit specific should be registered by calling register_trans_unit_variable. See the Multiple Translation Units chapter.

3.3.2. Include Files#

Include files always include any other files on which they depend. They also define a flag once they have been included, and will skip themselves if included again. (For example, cmd_line.h defines CMD_LINE_H). Therefore, one can include the files in any order and without regard for whether or not they have been included by other include files. To improve compilation speed, #includes can be surrounded by #ifndef flag/#endif.

3.3.3. Debug Code#

Debug code is surrounded by #if DEBUG/#endif. In such code, one should test the value of debug_level to control output of debug information. debug_level is zero ordinarily, to indicate that no debug output should be produced. Debug code should be set to come out at an appropriate level: 1 is for output that comes out once per source line, 3 for output that comes out once per token, and 5 for output one would hardly ever want to see. 2 and 4 lie in between those levels in the expected way. Choosing an appropriate debug level is more of an art than a science. If in doubt, make a guess at the best level, run the front end with that level of debug output, and see if the new messages seem out of place in the overall debug output at that level.

One can also use db_flag_is_set to determine whether specific output should be generated. For example, calls of db_flag_is_set("vtbl") are used in IL lowering to control debug output related to virtual function tables. By specifying -d-vtbl on the command line, one can get that particular debug output and nothing else. A particular “flag” is considered set if the string of the db_flag_is_set call is specified in a command-line debug option, so a new “flag” is added by adding a call of db_flag_is_set with some new string; there is no central registry of debug “flags.”

A similar facility for tracing names is provided by the --db_name=xxxx command-line option and the db_has_traced_name macro. This can be used to control extra output about processing of entities with specific names. db_trace provides a combination of the features of db_flag_is_set and db_has_traced_name.

Debugging options that are specified with the -d command-line option can also be specified using the db_opt pragma. The pragma can be used to specify a certain debug level or to define a certain debug flag for a portion of a source program that is being compiled. For example:

template <class T> void f() {
#pragma db_opt -nondep_call
  g();
#pragma db_opt #nondep_call
}

When a flag name is preceded by “#”, the flag is cleared if previously set.

Debugging options that are specified with the --db_name command-line option can also be specified using the db_name pragma.

The variable db_active is TRUE if debug_level is non-zero or if there is the potential for it becoming non-zero (i.e., if there is a list of debug requests from the command line). This can be used to control processing that may be needed when debug output is generated, but should not be done when no debugging is requested.

Diagnostic messages should be written to f_error. This is normally set to stderr, but in some modes is set to stdout, and can be redirected to a file by use of a command-line option. Debug output should be written to f_debug. This is set to f_error, but is a separate file variable to allow it to be changed if desired.

Entries and exits of routines should have calls to db_enter and db_exit. These provide flow tracing. The db_enter call looks like

db_enter(level, "routine-name");

The db_exit call has no arguments. These calls will cause an entry message and an exit message when debug_level is equal to or greater than the level specified. Note that the calls need not be surrounded by #if DEBUG/#endif; db_enter and db_exit are actually macros, and when DEBUG is 0, they are defined as null strings. When DEBUG is non-zero, they expand to calls to debug_enter and debug_exit. Those calls are surrounded by ifs testing db_active so that the routine calls are only done when debug output is actually being generated (or may be generated, when a list of routine-related debug requests is specified on the command line).

It should be clear that for db_exit to work correctly one should have a single exit point for each routine, and not have return statements in the middle of routines.

Some routines do not have db_enter and db_exit calls for speed reasons – even the overhead of two ifs would be visible. An example is skip_white_space in lexical.c.

3.3.4. Checking Code#

Checking code performs dynamic consistency checking, e.g., verifying the entry conditions for routines and the consistency of tables. Checking code is always surrounded by #if CHECKING/#endif. When a problem is found, this code calls

internal_error("routine-name: problem found");

which displays the argument string and causes abnormal termination of the front end.

As a concrete example of checking code: all switch statements should have a default clause; if the default clause does not correspond to any legal possibility, it should be present as checking code, and should call internal_error to indicate that a bad value appeared as the switch selector. (See CHECK_SWITCH_DEFAULT_UNEXPECTED in basics.h and default_is_unexpected in checking.h for macros that assist in applying this rule.)

Several routines are provided as convenient ways to do consistency checking. Calls of these routines need not be surrounded by #if CHECKING/#endif; they are defined as macros that expand to nothing when CHECKING is disabled. check_assertion can be used to check an expression to make sure it is true. expect_error checks that total_errors is nonzero at the point of invocation or that it becomes nonzero before the end of the compilation. This is useful to ensure that situations that are expected to lead to error diagnostics will indeed do so (only the first invocation will be reported on failure). check_assertion_or_expect_error combines the previous two routines: it has no effect if either the given condition is TRUE, or if total_errors is nonzero by the end of the compilation. unexpected_condition is used when a piece of code is not expected to be reached (for example, the default clause of a switch). Three versions of each of these routines are provided: the version described above, a version with a _str suffix that accepts a single string argument, and a version with a _str2 suffix that accepts two strings. The string(s) are printed on error termination to provide additional information about the cause of the error.

The consistency checks that are enabled by the CHECKING flag do not significantly affect the performance of the front end. There are other consistency checks whose execution does affect the performance of the front end. These extra consistency checks are enabled by the EXPENSIVE_CHECKING flag. These checks are not intended to be enabled when optimum performance is required (e.g., in production versions of a product).

3.3.5. Conventions for Declarations#

The type a_boolean (defined by default as int) should be used for true/false variables, along with the constants TRUE and FALSE. These are defined in basics.h and thus are available everywhere. In structs, a_byte_boolean is used for boolean fields when there is just one of them. When there are several, bit fields of length one are used, to save space. Obviously, these rules will not work on all host machines, but they will work well enough on most.

Fields in structs are ordered to some extent to reduce the space wasted in gaps for alignment. It’s assumed that pointers and large integers have similar size and alignment requirements, and thus a sequence of them will leave no gaps. When small fields appear, however, they are grouped together to avoid the gaps one would expect from a series of short/long/short/long fields.

Bit fields are declared with the type a_bit_field. It is usually unsigned int, but when compiling with a compiler that uses the Microsoft bit-field layout algorithm it can be set to unsigned char to get better packing of the bit fields.

The type a_byte should be used for small integer fields (i.e., fields whose values lie in the range 0-127), rather than char or unsigned char. enum types are often given a base type of a_byte so that their size is explicitly controlled to save space in structs.

For each table type (whether or not it is part of the intermediate language) there is an associated allocation routine and an associated routine that will set the fields of the table to default values. (In some cases, these two functions are provided by one routine.) When a table has a variant section and an associated tag field, there is also a routine to set the tag field and set the associated variant fields to default values. A tag field should never be assigned directly; it should only be set by this routine. Consistent use of the allocation, clear, and set-kind routines ensures that no uninitialized fields will cause sporadic problems, and that new fields can be easily added. In addition, it can be useful to add the following snippet to defines.h :

#if UNION_AS_STRUCT
#define union struct
#endif

Setting UNION_AS_STRUCT in a testing build allocates each variant field in its own storage. This often allows detection of cases where code accesses an incorrect variant for a given IL entry.

The names of macros are written in upper case in the front end source when they are object-like (e.g., manifest constants), but in lower case when they are function-like. The names of types and enumerators are written in lower case.

The names of typedefed, class, and templated types always begin with a_ or an_ (e.g., a_type, an_expr_node). Pointer type names end in _ptr. The names of constants in enumerations always begin with a prefix that hints at the enumeration type (e.g., enk_ for an_expr_node_kind).

3.3.6. Storage Allocation#

Allocation of space is never done by calling malloc or the like directly. For all table entries, as mentioned earlier, the associated allocation routine should be called. It will allocate and clear the space, and will also make sure that the space is allocated in the right area (e.g., intermediate language constructs must be allocated in the right memory region). For any other allocation cases, one of the routines in mem_manage.c should be called instead. alloc_fe has an interface like malloc, and allocates storage that need not survive the execution of the front end. alloc_general is similar, and allocates storage that will survive into the back end if the back end is called as part of the same program (this should not be used ordinarily).

One case to do carefully: all names that become part of the intermediate language (file names, directory names, identifier names) should be allocated in the file-scope intermediate language region, by calling alloc_il.

Space is never freed by calling free or related routines. Where tables are reused, the front end maintains individual lists of available entries for each type of table, and takes entries off those lists before allocating new ones. Space can be released on a memory region basis, for example to free up all the space used for a given function once the function has been processed. When the front end is configured in such a way that it can be re-initialized (by setting MAKE_FRONT_END_CALLABLE) the front end keeps a record of all memory allocations and automatically frees the memory when the front end terminates.

In general, every reasonable attempt is made to not allocate tables unless they are actually needed. The intermediate language, for example, will contain some wasted space, but extremely little.

In order to avoid any fixed limits in the front end, the primary non-IL data structures are designed so they can be dynamically allocated and expanded (curr_source_line, scope_stack, etc.). An initial allocation is done that should cover all expected uses; if more space is needed, the table is reallocated with a larger size. That means that there is a danger inherent in keeping around pointers into those data structures, because if the table moves, all pointers to the table become obsolete. In the input line/macro area, such pointers are necessary, and there is a pointer registration scheme that is used to make sure that all pointers are updated if a reallocation occurs. In other areas, pointers can be avoided relatively easily. A simple rule to remember is that one should never make a local pointer into one of the primary stacks (scope_stack, struct_stmt_stack); one should instead use an index value into the stack, since an index is not invalidated by a reallocation of a table. When memory is allocated that may later be resized, it must initially be allocated using either alloc_resizable_buffer or realloc_buffer. It may then be resized later by calling realloc_buffer.

One can get a summary of space use by having debug_level non-zero in fe_wrapup. This can be requested by specifying -dfe_wrapup=1 on the command line. A summary of symbol table and constant sharing efficiency is also included.

3.3.7. Typerefs#

The description of a type in the intermediate language contains a type entry kind tk_typeref, which is used to define typedefs, and also to add type qualifiers to a type, e.g., const in const int. Such an entry is in some ways not a full-fledged type. For example, it doesn’t indicate its size; one must go to the underlying type to get that. The macro skip_typerefs and the function f_skip_typerefs can be used to strip off typerefs and get to the underlying type, which can then be processed in the normal way. A very common cause of bugs in code our customers add is failure to put in skip_typerefs calls where necessary. (For that matter, it’s a common cause of bugs in our code, though we try hard to find all those bugs before we ship the code.) The whole skip_typerefs issue is error-prone and requires discipline, but you do get used to it, and after many years of working with it we remain convinced that it’s the right way to handle typedefs and type qualifiers. Just be careful.

3.3.8. Source Positions#

Source positions are represented as a structure of type a_source_position (see basics.h). It contains a sequence number and a column number. The sequence number is a unique number assigned in ascending order to each source line as it is read. Unlike a file name and line number, the sequence number is unambiguous. It’s also more compact, and can be converted to a file name and line number for external use (this is done in error messages, for example).

The column number is a 1-origined character position in the original source line. The character position is the logical character number in the line, not simply a byte offset. This is significant if the source file contains multibyte characters or is in a format such as UTF-16. Tab characters are counted as a single column. The position of text within a macro expansion is considered to be the first character of the macro name in the macro call in the original source line that directly or indirectly generated the text.

If the front end is configured with FULLY_RESOLVED_MACRO_POSITIONS set to TRUE, each source position contains an additional sequence/column number pair. For text within a macro expansion, this additional position information gives the location from which the associated text was copied into the macro expansion, i.e., a macro definition or macro argument.

If MACRO_INVOCATION_TREE_IN_IL is configured to TRUE, each source position referring to text within a macro expansion contains the index of the macro invocation record that identifies the macro in whose expansion the text at that position occurs. For text appearing directly in the source, this index will be NO_PARENT_MACRO_INVOCATION.

A sequence number of zero is used to indicate an unknown position or a “position” in initialization, in a predefined macro, or in the command-line. The column number indicates which. See basics.h.

Importing assembly metadata in C++/CLI is a process that creates a string containing C++ code that expresses that metadata and then parses the contents of that string. Each import is associated with a particular sequence number, and that number is assigned to the position of each token extracted from that string (with a column number zero). For example, when compiling a C++/CLI source file, sequence number one will usually correspond to mscorlib.dll.

3.3.9. Language Dialect#

The global variable C_dialect (in cmd_line.h) indicates the C dialect to be accepted: the choices are C_dialect_cplusplus, C_dialect_ANSI, and C_dialect_pcc. The macro C_mode() is useful as a way of testing for C versus C++. cppNN_mode (where NN is the year of the relevant C++ Standard, e.g., cpp03_mode, cpp17_mode, etc.) is TRUE when the features of the corresponding revision of the language should be accepted, and similarly cNN_mode for versions of the C language.

The global variable strict_ansi_mode is TRUE if use of features that are not part of the selected C or C++ Standard should cause diagnostics. strict_ansi_error_severity indicates the level of diagnostic (warning, discretionary error, or error) that should be produced for such cases.

microsoft_mode indicates whether a Microsoft dialect (C, C++, C++/CLI, or C++/CX) is in effect. In configurations that do not allow Microsoft extensions (i.e., when MICROSOFT_EXTENSIONS_ALLOWED is FALSE), it is a macro equivalent to FALSE. This allows many Microsoft-mode-specific bits of code to be written without surrounding them with conditional compilation directives testing MICROSOFT_EXTENSIONS_ALLOWED. cppcli_enabled and cppcx_enabled similarly indicate that the Microsoft dialect is C++/CLI or C++/CX, respectively (if TRUE, microsoft_mode is TRUE too, and C_mode() is FALSE). Some C++/CLI- and C++/CX-specific code is delimited by conditional compilation directives testing MICROSOFT_EXTENSIONS_ALLOWED because it uses variables, data members, or functions that are available only when that macro is TRUE. (The macro CPPCLI_ENABLING_POSSIBLE is tested in a few places for code that deals with the importing of assembly metadata, but it shouldn’t be used to conditionalize other C++/CLI-specific code. In particular, the IL extensions representing C++/CLI and C++/CX constructs are only conditioned on MICROSOFT_EXTENSIONS_ALLOWED.)

gpp_mode and gcc_mode indicate that, respectively, GNU C++ or GNU C emulation is in effect. In configurations where GNU_EXTENSIONS_ALLOWED is FALSE, these are macros equivalent to FALSE.

clang_mode indicates that clang emulation is in effect. Because the initial emulation of the clang C++ compiler in the front end was very close to g++, clang_mode also implies that gpp_mode or gcc_mode will be set. When it is necessary to distinguish between the GNU and clang emulations, the macros gnu_version_is, gpp_version_is, gcc_version_is, clang_version_is, clangcpp_version_is, and clangc_version_is (see lang_feat.h) allow finer-grained discrimination (including the compiler version to be emulated). (There are also corresponding Microsoft-specific macros, ms_version_is, mscpp_version_is, and msc_version_is.)

cfront_2_1_mode is TRUE if the cfront 2.1 dialect of C++ is to be compiled. cfront_3_0_mode indicates the cfront 3.0 dialect. The macro any_cfront_mode() tests for either of those being set to TRUE.

(There are other variables/macros that control dialect variations. The ones mentioned here are among the more frequently used ones.)

3.3.10. Expression Processing#

See Rescanning Expressions for some additional coding rules that apply only within the expression-processing routines.

3.4. Recursive Descent Parsing#

“Recursive descent” is an ad-hoc syntax parsing technique. It works well for languages (like C) where very little lookahead is required to resolve local ambiguities, and where the language in which the compiler is written is recursive (as is true for the front end).

The idea of recursive descent provides a framework and a discipline for parsing a language, but it is implemented entirely by hand-crafted code. There is nothing automatic about the parsing. Thus, there is more possibility for error in developing a recursive-descent parser than there is in developing a table-driven parser. However, if one writes carefully and tests thoroughly, the end result is better with a recursive descent parser: faster parsing, better error recovery, and a clearer compiler (because all of the logic related to a particular construct appears in one place).

To implement recursive descent, one writes routines for each non-terminal construct in the language syntax. For example, one writes a routine to scan an expression, and one to scan a statement, and one to scan a declaration. The internal structure of these routines looks a great deal like the syntax of the constructs they implement. For example, if a declaration is defined as a list of declaration specifiers followed by a declarator followed by a semicolon, the routine that scans a declaration would call the routine to scan a list of declaration specifiers, then the routine to scan a declarator, and then check for and discard a semicolon. Where the language has recursive constructs, the routines involved will call each other recursively.

The terminal tokens that are being parsed come from the lexical routines, specifically from get_token. At any given moment, there is a current token (curr_token). It is the next token of input, the one that hasn’t yet been taken. If the token is an identifier or a literal constant, there is additional information associated with it. See lexical.h.

As the recursive descent routines parse the input tokens, they build up information on what they have parsed, in the intermediate language trees and in the symbol table. They also check for and recover from errors.

The scanning of expressions is one area where recursive descent is inefficient for languages (like C) where the expression syntax has many levels. A naive recursive descent implementation would provide a routine for each level of the expression syntax, and thus scanning a simple expression would involve perhaps a dozen subroutine calls. For this reason, the front end uses a modified recursive descent technique to scan expressions. Internal calls to the scan_expr routine (in expr.c) specify a precedence level, and the scan routine can decide whether to call itself recursively or not based on the relative precedences of adjacent operators. With this modification, the number of calls is reduced to very nearly one per operator. See the description of expr.c later in this document for more detail.

To reduce the possibilities for errors in writing the parsing routines, one should not check for the start of constructs by checking directly for all the possible first tokens in that construct. Such a list of tokens is inherently error-prone. For constructs of this kind, there are functions that will indicate whether or not the current token looks like the start of the construct (see, for example, is_decl_start in decls.c). Such routines might themselves call other is_xxx_start routines if the constructs can themselves begin with other complex constructs.

Several utility routines in lexical.c are useful for recursive descent parsing. required_token can be used to check for and discard a required token, and to issue a specific error message if the token is not present. loop_token is used to check (at the bottom of the loop implementing an iterative syntactic construct) for the occurrence of a particular token. The token is taken if found, but no error is given if the token is not found. next_token is useful for doing a limited kind of lookahead: it will return the kind of the token following the current token, without actually getting that token (actually, it gets it and then puts it back). In many cases, that one token is enough to allow a decision to be made at a fork in the syntax. See, for example, the code that recognizes statement labels in statement in statements.c.

Consult any of the popular books on compiler theory for a more detailed overview of recursive descent.

3.5. Error Recovery#

When the front end discovers an error, it has several goals:

  • To describe the error clearly and indicate its position (including the position within the source line);
  • to recover from the error without generating extra errors;
  • to preserve as much information as possible about the valid parts of the construct containing the error, so that detection of other errors is not needlessly compromised.

(“Error” here of course refers to remarks and warnings as well.)

Describing an error clearly requires that one not only detect the error, but detect it at an appropriate level, so that the front end can guess better at what was intended. Sometimes this requires clever code at decision points in recursive descent parsing, so that on invalid constructs one goes to the more likely of two subcases. Within the { } for a struct, for example, we have a series of member declarations separated by semicolons. If a semicolon is missing, one should not fall out of the loop and leave the scanning of the struct. In other situations, clearly invalid cases should be flagged at the higher level, and none of the subcases should be entered. For example, at the beginning of a statement or a declaration, if the first token is not a valid start for that construct, an error message (“expected a statement” or “expected a declaration”) is more informative than descending into any of the subcases and letting them issue errors.

3.5.1. Error Positions#

The position of an error is often just as important as the error text. The front end tries to give useful error positions. This is aided by the fact that column positions are maintained for each token. A global variable error_position indicates the position of the beginning of the last language construct scanned. It is set by get_token on scanning each token, and it is also set by each medium-sized language construct (e.g., an expression, a declarator). Most recursive descent routines save the position of the first token of the construct they scan, and then set error_position to that position on return. error_position is the default position for errors when no explicit position is specified, and this protocol for setting it works well in ensuring that the error positions are right. Some cases take extra work. For example, in the expression routines, the position of the left operand and right operand, and the position of the operator itself, are known for each operation. If an error is due to a problem with one of the operands, the error position used is the position of that operand. If the problem is with the combination of those operands (as in adding a pointer and a float), the error position used is the position of the operator.

The macro set_err_pos_to_curr_token is used when one is going forward in scanning. It sets the error position to the position of the current token. One would use it, for example, when looking for the = that indicates an initializer in a declaration. The previous construct (a declarator) has been scanned and checked, and its position is in error_position. Now, however, we are going forward, and the = that may have ended the expression scan should once again be the current construct. In practice, there are few cases like this; they only come up for optional tokens in the middle of constructs. required_token handles the change of error position in the more usual case of a required token.

When an error refers back to something earlier in the source, it often makes sense to give the position of the other location rather than the current location. Variables declared but never used, for example (a warning), are detected at the closing brace of functions. However, that position is much less informative than the position of the declaration of the variable in question, so that is used. Likewise, when a comment is unclosed at end of file, it is more useful to indicate the position of the start of the comment instead of the location of the end of file, where the error was detected.

3.5.2. Error Codes#

Error messages are mapped to values of the enumeration an_error_code (see err_codes.h, along with err_data.h for the associated error message text). For example, ec_exp_line_number represents the message “expected a line number”. ec_no_error is first in the enumeration, and therefore has the value zero. Thanks to it, one can return error codes of type an_error_code, and test them in if statements (non-zero means an error).

Errors with codes beginning ec_exp_ (and texts beginning “expected a …”) are uniformly syntax errors – some expected token or construct was not found.

The severity of an error is determined by the error routine called (e.g., error instead of warning) and not by the error code. In fact, there are cases where the same error code is used as both a warning and an error. The severity of certain diagnostic messages may be overridden using a command line option. A diagnostic that is issued with a severity that is less than an error (i.e., remark, warning, or discretionary error) may have its severity changed, or may be suppressed completely.

Some error message texts have fill-ins, as in

could not open source file "name"

For these, variants of the error routines should be called. For the case above, for example, str_error would be called. The fill-in string is passed to the error routine as a null-terminated string, and replaces %s where it appears in the error text string. See Error Reporting for details on the kinds of fill-ins available.

3.5.3. Syntax Errors#

In terms of internal processing, one must distinguish syntax errors from semantic errors. Syntax errors are discovered when the sequence of tokens in the source program cannot be parsed according to the language syntax (for example, i=j+;–there’s something missing between the + and the ;). All other errors are semantic errors – things that can be parsed but are still nonsensical (for example, use of an undefined variable, or adding two structs). In error recovery, the important distinction between the two is that for a syntax error one must decide how to continue parsing, whereas for a semantic error everything is fine as far as parsing and there’s no question about how to continue.

Syntax errors should always cause the calling of flush_tokens in lexical.c, and semantic errors should never call it. The routine syntax_error records the occurrence of an error, and then calls flush_tokens, which is a convenient combination.

flush_tokens scans and throws again tokens until it finds one in the stop tokens set, and then returns. Paired tokens are also considered: if a ( is scanned while flushing tokens, for example, the corresponding ) will be sought before continuing the search for a stop token. There are some other special cases; see the code in lexical.c.

The stop tokens set is an array indexed by token kind. Each element of the array is non-zero if the corresponding token is currently in the stop tokens set (and thus should halt a flush). As recursive descent routines scan constructs, they add to the stop tokens set those terminal tokens they expect to see in their constructs, and then remove those tokens once they are encountered. The macros add_stop_token and remove_stop_token are called to add and remove tokens from the set by incrementing and decrementing the associated element of the array. Since syntactic constructs are nested inside one another, having each recursive descent routine add the tokens it recognizes produces a stop tokens set that will halt a flush at a token with which one of the active routines feels it can continue.

At boundaries where the syntax changes dramatically, push_stop_token_stack is called to establish a new set of stop token values. The old values are restored later by calling pop_stop_token_stack. This is done, for example, when entering the processing for preprocessing directives, whose syntax is dramatically different than the context in which they are embedded.

Uses of add_stop_token and remove_stop_token must be very carefully paired. A missing call of either would probably cause the associated token to be handled incorrectly in error flushes for the rest of the compilation. Skipping a remove_stop_token call is easy to do if one does a goto out of an inner loop in a recursive descent routine (say, because there is an error) and one skips over the point where the remove_stop_token is done. The likelihood of such a problem is lessened by having debug_enter and debug_exit verify (by means of a checksum) that the contents of stop_token_array are the same on exit as they were on entry. This check is only done when debugging is active, however, because it’s time-consuming. Also, fe_wrapup verifies that all the counts have returned to zero by the end of the compilation; that check is always done. In addition, when expensive checking code is enabled, pop_stop_token_stack verifies that all of the elements of the stop token set that is being discarded have been reset to zero.

If required_token does not find the token it expects, it adds that token to the stop tokens set, and calls syntax_error. After flush_tokens finishes, required_token checks again to see if the desired token has come up, and if so, scans over it.

3.5.4. Error Entries#

For both syntax and semantic errors, another part of error recovery is producing something sensible, in spite of the error, in the data structure that represents the program. This task is eased by having special “error” entries for all of the major table types. There are, for example, error constants, error types, and error expression nodes. They are used to take the place of any part of the program tree that cannot be built because of an error. In general, the code tries to build as much of the complete tree as possible. An error in one part of an expression, for example, need not mean that the entire expression tree will be replaced by an error expression node. Throughout the front end, code must be capable of dealing with error entries. The basic philosophy is that an error entry is considered to be compatible with anything else. No additional errors should be issued on encountering an error entry, since it is assumed that some error was issued when the entry was created.

When syntax errors involve missing identifiers, special error symbols are created. A different error symbol is created for each error case. All error symbols point to the same special associated symbol header (see the documentation of symbol_tbl.c).

When errors involve duplicate declarations of the same symbol, the second symbol is entered anyway. Both symbols are in the symbol table, but because of the symbol table structure, the later declaration will the one that is visible, and it therefore effectively replaces the earlier declaration.

In general, scanning routines do not need to return indications of success or failure, since they always return a tree representing the construct scanned. Some routines must deal with unusual cases where there is no possible representation for the thing scanned, and so they return an error parameter. The caller must check it.

3.5.5. Intelligent Error Diagnosis#

The front end attempts to diagnose as many errors and warnings as possible. If one uses header files with function prototypes for all external functions and variables, it should be possible when writing in ANSI C or in C++ to do away with lint, or at least to relegate it to use in unusual cases. That is the goal of the front end: diagnose not only errors but also questionable constructs such as those flagged by lint. However, since lint is run on request and the front end must be run on each compilation, one must take care that the errors are not overly pedantic. The front end tries to do this by being selective about errors. It’s a good idea, for example, to warn about value-less return statements in functions that should return a value. However, in old-style C, before the keyword void existed, it was standard practice to omit the type specifier altogether on a function that returned nothing. The C language says such a function has type int, but the C programmer usually thinks of such a function as having type void. So, in the front end, a warning is issued for a value-less return in a function with an explicit type, but not for one in a function with an implicit int type. In pcc mode, a warning is never issued.

A similar issue comes up in expressions. An expression like 1/0 looks like an error (division by zero), but should not necessarily cause a compilation error. If it occurs in a context where it is not evaluated (such as 0&&1/0), it never causes an error; if it occurs in a context where it must be evaluated at compile time (as in the size of an array), it causes an error; and if it occurs elsewhere (i.e., in executable code), it causes a warning and is left as-is to cause a fault (or not) at execution time.

Likewise, when a constant seems to be part of a bit-oriented operation (it is a non-decimal constant, or it results from a bit-wise operator applied to constant operands), the normal warnings about implicit sign changes when the constant is changed from signed to unsigned or vice-versa are suppressed.