18. Administrative Functions#

The source file cfe.c contains the main program. It calls proc_command_line (in cmd_line.c) to parse the command line; fe_init (in fe_init.c) to do global initialization; translation_unit (in decls.c) to compile or cpp_driver (in preproc.c) to do preprocessing; and fe_wrapup (in fe_wrapup.c) to do global termination. If a back end should be called, and there were no errors, it also calls back_end. Then, it calls fe_wrapup_part_2 (also in fe_wrapup.c) to do post-back-end wrapup. Finally, it calls exit_compilation to exit with the exit status appropriate to the severity of the errors detected.

If multiple source files can be compiled, the main program includes a loop around the steps from fe_init to back_end. It calls get_next_source_file to see if there are any more source file names in the command line.

Usually, the front end has its own main program, but it can instead be used as part of some other main program. This is done by setting the MAKE_FRONT_END_CALLABLE configuration macro. The name of the main entry point into the front end is specified by the EDG_MAIN macro, and is edg_main by default. When the front end is callable, exit_compilation does not actually exit the program, but instead returns control to the routine that called the front end.

18.1. Initialization and Wrapup#

fe_init.c contains one-time initialization. Declarations of external variables in .h files, which must begin with EXTERN, are compiled here with EXTERN defined as an empty string, so storage is actually allocated. Likewise, VAR_INITIALIZERS is defined as 1, so that initializers on those variables are put out. The main routine is fe_init; it calls other routines to initialize predefined macros, keywords, and assorted state variables. In particular, it calls some initialization routines in other files (e.g., lexical_init, sym_tbl_init). There are several different kinds of initialization routines:

  • “early” initialization routines (e.g., cmd_line_early_init), which are called before command-line processing is done.
  • “one time” initialization routines (e.g., lexical_one_time_init), which are called only once per invocation of the front end, even if multiple source files have been specified on the command-line (when COMPILE_MULTIPLE_SOURCE_FILES is enabled).
  • “per-compilation” initialization routines (e.g., lexical_init), which are called for each source file specified on the command-line (when COMPILE_MULTIPLE_SOURCE_FILES is enabled).
  • “trans-unit” initialization routines (e.g., lexical_trans_unit_init), which are called once for each translation unit processed when processing exported template definition files or when COMPILE_MULTIPLE_TRANSLATION_UNITS is enabled.

Because the front end can be reinitialized (when MAKE_FRONT_END_CALLABLE is used), most variables must be initialized dynamically, not statically. Variables that are modified by command-line processing must be initialized in early initialization routines.

Command-line -D and -U options are also handled here (see preproc_init).

If multiple source files are accepted, fe_init is called at the start of compilation of each file.

fe_wrapup.c contains one-time wrapup code. This involves popping the file name scope, checking for errors on files written, and dumping debug statistics, such as information on symbol table efficiency and on total memory used. fe_wrapup is called after the front end is done but before the back end (if any) is executed, and fe_wrapup_part_2 is called after the back end (if any) is executed.

18.1.1. C++/CLI Initialization#

The C++/CLI language depends in some essential ways on its core library, which is loaded from assembly metadata (usually a file called mscorlib.dll). The initialization process therefore preloads that library and any others specified via the --preusing command-line option with a call to process_preusings.

Immediately after loading the top-level declarations of the core library, process_preusings also initializes a number of global variables (and global array elements) to point to the symbols for special entities from the core library (for example, the ref class System::Object). This is done in init_cli_symbols.

init_cli_symbols also creates templates for cli::array, cli::interior_ptr, and cli::pin_ptr. This is achieved by running ordinary front end processing on string literals (via calls to scan_top_level_metadata_declarations) containing the declarations for those templates. The symbols representing these templates are recorded at this time (for ease of reference in later contexts).

The symbol for namespace ::cli itself is created even earlier by a call to make_symbol_for_namespace_cli.

18.2. Debug Routines#

The source file debug.c contains the routine that processes the command-line debug request option (-d), and also the routines debug_enter and debug_exit which are called for flow tracing. When the command-line debug request includes routine names, debug_enter checks the name of the routine entered against the names on the request list. If the name is found, the associated modification of debug_level is done. An internal stack is used to preserve information from debug_enter time to debug_exit time for a function. One trick: A checksum of the stop tokens set is computed and stored on entry, and recomputed and checked on exit. This helps catch mismatched add_stop_token and remove_stop_token calls in recursive descent scanning routines.

debug_enter and debug_exit are not called if there is no command-line debug request (i.e., if db_active is FALSE).

18.3. Generating Strings from IL Types and Constants#

il_to_str.c contains code that converts IL types and constants into C/C++-form strings. Associated declarations are in il_to_str.h.

To use these routines, one needs to declare a control block of type an_il_to_str_output_control_block, call clear_il_to_str_output_control_block to initialize it, and then set applicable fields to describe how the string output should be done. The control block includes, among other things, a pointer to a routine to output a string, so one can establish a control block that provides an output routine that writes to a file, or a control block that provides a routine that puts the output string into a buffer, etc. Other fields in the control block control the style of output and provide other (optional) callback routines to do specialized output (instead of the default version) for certain constructs.

form_type produces a string for a type. It does this by calling

  • form_type_first_part, which outputs the type specifiers and the part of the declarator that precedes the name, and then
  • form_type_second_part, which outputs the part of the declarator following the name.

So, for example, for the type

array [10] of pointer to const int

the “first part” would be

const int *

and the second part would be

[10]

If one wants to put out a declaration of an identifier, one can do so by

  1. calling form_type_first_part, then

  2. outputting the identifier, then

  3. calling form_type_second_part.

producing an overall output string like “const int *a[10]”.

form_type_first_part descends through any derived types that are at the top of the type by calling itself recursively. When it reaches a non-derived type, it calls form_type_specifier to output the specifiers part of the type (e.g., “const int” for the above example). Then, while reascending, it puts out the parts of the declarator that appear on the left-hand side of the declarator (pointers, references, pointers-to-members, and any required parentheses).

form_type_second_part also descends through any derived types by calling itself recursively. While descending, it puts out the parts of the declarator that appear on the right-hand side of the declarator (arrays, functions, and any required parentheses). form_array_declarator puts out array declarators. form_function_declarator puts out function declarators; it can be replaced by a user-provided routine if desired. One reason to replace it: It does not handle putting out function declarators for function definitions.

form_type_specifier puts out non-derived types. Most cases are simple (see, for example, form_int_kind_name and form_float_kind_name). Typedef names are put out as such. class, struct, union, and enum types are always put out as references to those types, i.e., a definition is never put out.

form_name puts out names; it can be replaced by a user-provided routine if desired. It handles qualified names (see form_class_or_namespace_qualifier), unnamed entities, and template argument lists on classes (see form_template_args).

form_constant produces a string for a constant. It does some processing itself, and has several important helper routines:

  • form_integer_constant puts out integer constants, including sign, suffix, and implicit casts. It also handles integer constants cast to pointer types (e.g., null pointers).
  • form_address_constant puts out address constants, including the “&” and optimizations for arrays and for adding a byte offset to the base address.
  • form_pm_constant puts out pointer-to-member constants, including null constants and any required implicit casts.
  • form_char puts out a single character of a string or character constant, dealing with unprintable characters and the like.

18.4. Error Reporting#

error.c contains code that handles reporting of errors. error.h contains the declarations needed to use the error routines. error_msg.txt defines the error codes and associated error message text for all of the error messages, and, along with error_tag.h, is used by mk_errinfo to generate the err_codes.h and err_data.h files. It is the generated files that are actually used by the front end to define the an_error_code enumeration and the array of error message text. See mk_errinfo for a description.

an_error_code defines enumeration constants that represent each error message (for example, ec_unclosed_string represents the error message “missing closing quote”). ec_no_error is the first member of the enumeration, and therefore has the value 0, which has the effect of allowing flags with the value 0 for no error, and a non-zero value to indicate a specific error.

Within the error message texts, fill-in codes beginning with “%” indicate points at which information is to be inserted:

%%
A “%” is inserted into the message.
%d
A number is inserted into the message.
%s
A character string is inserted into the message. %sq indicates that the inserted string should be placed in quotes, e.g.,

could not open source file "abc.h"

%t
A description of a type is inserted into the message. It is always placed in quotes, e.g.,

cast to array type is nonstandard (treated as cast to "int *")

%n
A description of an entity in the program (identified by a symbol) is inserted into the message. The default %n (“name”) can be modified by appending an “o” (%no, “name only”) to get just the name of the entity, a “p” (%np, “name with function parameters”) to force parameter information on a function, an “f” (%nf, “name full”) to get full type information on the name, an “a” (%na, “name with template arguments”) to get arguments on a template, “t” to force the “A<T> with T=int” form of template name to be used, or a “T” to get the translation unit in which the symbol was declared.
%n
function "A::f"
%no
"A::f"
%np
"A::f(int, float)"
%na
"A<int>"
%nt
"A<T> with T=int"
%nT
function "f(int)" (from translation unit "x.c")
%p
A source position is inserted into the message.
%[label]
The specified label is looked up in the label_fill_ins array and one of two error messages is substituted depending on the value of a boolean variable associated with the label. The substituted error messages cannot themselves have fill-ins. See the definition of label_fill_ins to determine the list of valid label fill-ins (or to add a new one).
%T
A template argument list is inserted into the message.

In the default mode, the entity description is preceded by a keyword that identifies the type of the entity (e.g., function above). That is also done in the “full” mode for entities that are not variable-like or function-like. The normal “%n” displays function parameter types only for overloaded functions. Any of the %n codes can be modified by appending a “d” to request information on the point of declaration of the symbol:

declaration is incompatible with "int f(char)" (declared at line 17)

Following the optional “d” modifier, the “T” modifier may be used to display the translation unit associated with the symbol (”from translation unit "x.c"”). When compiling a primary translation unit, this will cause the translation unit to be displayed for symbols from secondary translation units. When compiling a secondary translation unit, this will cause the translation unit to be displayed for all symbols.

There can be more than one fill-in code in a message. When there is, each fill-in code can be numbered (e.g., %t1, %t2) to tie it to a corresponding fill-in argument. Note that this technique allows one to reorder the fill-ins in a message (say, when changing the error messages to a language other than English) without changing the order of the fill-in arguments on error calls throughout the front end.

Each normal diagnostic includes four kinds of information:

  • the error code, which is a value from the enumeration just mentioned, and indicates the error message text;
  • a severity (remark, warning, discretionary error, error, or catastrophic error);
  • the error position (a sequence number and column); and
  • optionally, some information that goes into fill-in spots in the error message text.

The severity is usually indicated by the routine called (e.g., warning instead of error). There are also routines that accept the severity as an argument. A given message can be (and is, in many cases) output with different severities in different modes or different places.

The severity of a given error code may be overridden through the use of a command line option or pragma. For example, the option --diag_error=extra_semicolon may be used to cause the extra semicolon diagnostic to be issued as an error, even if the warning function was actually called to report the error. Only diagnostics that would by default be issued as remarks, warnings, or discretionary errors may have their severities overridden (i.e., nondiscretionary errors and catastrophic errors may not have their severities overridden). When a severity is overridden by a pragma, the default severity can be restored by use of the diag_default pragma. The default severity is the severity as would normally be used for the message as adjusted by any command-line options.

The error severity may be overridden by specifying either an error “tag” or an error number. The error tag is a name assigned to each error code. The error number is the numeric value of the error code enumerator that represents the message. By default, the error tag is the same as the error code enumerator name with the ec_ prefix removed. The error tag may be changed by specifying an alternate tag in the error_msg.txt file. Additional tags may be provided by creating entries in the error_tag.txt file. set_severity_for_error_tag and set_severity_for_error_number are used to override the severity for a given diagnostic. These routines are called to process the error severity options during command line processing.

The error position is usually given by the global variable error_position, but it can also be explicitly supplied. error_position is set to the start of each token as it is scanned, and is set on exit from recursive descent routines to the start of the entity just scanned, unless that entity is extremely large (like several lines). This usually results in the correct position in error_position, at least for syntax errors.

When the sequence number of a_source_position is 0, the column has special meaning (for example, it can indicate that the error occurred in the command line). See basics.h.

The error routines themselves have names that encode the severity (one of remark, warning, error, catastrophe or the generic diagnostic), and the particular combination of fill-ins and a source position that they accept. That is, an error routine name may be as simple as error or as complicated as pos_ty2_warning. pos indicates that a source position is provided; str or st indicates that a string fill-in is provided; sym or sy indicates that a symbol fill-in is provided; type or ty indicates that a type fill-in is provided. These can be followed by a number if more than one fill-in is needed, e.g., ty2 for two types. So, for example, pos_ty2_warning issues a warning, and the caller is expected to provide the error code, the source position, and two fill-in types. See error.c for the full list of routines. Not all combinations exist, but it is easy to add any routine that does not exist.

For a syntax error, one can instead call

syntax_error(error_code);

which will call error and then flush_tokens (see lexical.c).

For complicated errors that may have to list an arbitrary number of items, for example diagnostics for overloading ambiguity, one can make an initial call to issue the initial line of the error (calling start_error or one of its variants such as type_start_error), then any number of calls to add individual lines of additional information (calling add_diag_info or one of its variants such as sym_add_diag_info), and then a call to end_error to finish the error processing (including writing the source line in error). The end result is illustrated by

"test.C", line 7: more than one constructor applies to convert from
          "double" to "A":
            function "A::A(int)"
            function "A::A(float)"
    f(2.0);
      ^

The usual fill-in processing is available on both the initial line and any lines of additional information. The additional lines are described by an error code, so they can also contain arbitrary text.

The error formatting and processing is done by construct_message. It fetches the error message text, builds a linked list of message segments describing the message and its fill-in codes, expands each fill-in appropriately, and writes the full message.

If an error is detected while processing a template instantiation or during implicit generation of a routine (e.g., a constructor), it is helpful to provide additional information on the context in which the error was detected:

"test.c", line 7: error: "A::A()" is inaccessible
  B x;
     ^
          detected during implicit generation of "B::B()" at line 7

The context is deduced from the entries in the scope stack. include_in_context_output identifies the scope stack entries that merit extra context information. If there are any such entries, diag_message outputs an extra context line for each, in order from innermost context to outermost.

format_output_line is the routine that actually prints the diagnostic messages; it is the routine to change to modify the error output. To output the source line associated with the diagnostic, write_source_line is called. If the error is in the current line, write_orig_source_line reconstructs and writes out the original source line, with a caret under the column position of the error. If the error is not in the current line, write_error_source_line reads the line from the source file, if possible, and writes out the source line. Long diagnostic messages are wrapped across multiple lines. If raw listing information is requested, similar information is written to f_raw_listing (in encoded form). The diagnostic routines count the number of diagnostics of each severity so that the compilation can be ended with an appropriate exit status. They suppress diagnostics that are below the error display threshold. If an error is catastrophic or if the error limit is reached, the compilation is terminated.

The severity of diagnostics for strict ANSI violations and C++ anachronisms can be specified by command line options. To simplify handling of these diagnostics the global variables strict_ansi_error_severity and strict_ansi_discretionary_severity contain the error severity to be used for strict ANSI violations. The former is either es_warning or es_error, while the latter is either es_warning or es_discretionary_error, depending on which strict mode was selected on the command line. anachronism_error_severity contains the error severity to be used when anachronistic features are used. These global variables can be used as arguments to the diagnostic reporting routines.

command_line_error and str_command_line_error are the routines called for command-line errors (without or with fill-in text, respectively). They do not return.

The routine internal_error should be called for internal errors, such as failures of consistency checks. It should not be called for anything that is due to a source program error or some environmental error (like a missing file). The call should look like

internal_error("routine-name: problem found");

Several routines are provided as convenient ways to do consistency checking. Calls of these routines need not be surrounded by #if CHECKING/#endif; they are defined as macros that expand to nothing when CHECKING is disabled. check_assertion can be used to check an expression to make sure it is true; check_assertion_str is similar, but includes a string to be printed on error termination. unexpected_condition is used when a piece of code is not expected to be reached (for example, the default clause of a switch); unexpected_condition_str is similar, but includes a string to be printed on error termination.

18.5. Memory Management#

mem_manage.c contains memory management routines. mem_manage.h contains the associated declarations.

Memory is allocated in memory regions. File-scope intermediate language information is put in memory region 1, and the functions in the source program will each have one of the memory regions 2, 3, and so forth. Memory region 0 is used for storage that need not survive the front end, like the symbol table.

Before using a memory region, one calls either init_memory_region (for regions 0 and 1) or new_memory_region (for regions for functions) to initialize the new region.

alloc_in_region is then called when needed to allocate space in a given region. (alloc_fe can be called for the special case of allocation in region 0, the front-end-only region.) alloc_in_region allocates a large block of memory for each memory region by calling malloc, then parcels out that storage as it is needed. When the block is exhausted, a new block is allocated and put on a linked list of blocks for that memory region.

When one is done with an IL memory region, check_for_done_with_memory_region is called; it will either write the information to a file and free the block, or trim the end of the memory region to reclaim any unused space (depending on whether an IL file is being generated), or defer this processing until later if the memory is potentially needed for operations such as inlining or generic lambda instantiation. “Trimming” means making a separate available block from the remaining space at the end of the last block in the region. Since the blocks are large (typically, 64K) and a typical function will use less than 10K of the block, this is an important feature. The combination of using a large block size and reusing the remaining space improves efficiency (in reducing the number of calls to malloc) and avoids the need to guess at the amount of storage a particular memory region is likely to need. Eventually, when the partial blocks created in this way are released, they are recombined into the original large blocks allocated via malloc, and freed using the system free. Thus storage used is freed, but only on a large-block basis, so fragmentation is not a problem.

alloc_general is a more-or-less direct interface to malloc to be used in the infrequent cases where storage must be allocated that will survive into the back end if the back end is called in the same program. An example of such storage is the debug option control list, which holds information about the command-line debug option.

alloc_resizable_buffer and realloc_buffer are used to allocate tables that can be enlarged (e.g., input_stack).

When MAKE_FRONT_END_CALLABLE is used, the front end keeps a record of all memory allocations so that all of the memory allocated can be freed before returning control to the caller.

When precompiled headers are used, the memory management routines work to allocate the IL memory region blocks at reproducible addresses, so that the information in a precompiled header file can be read back in without the need to walk the IL and remap pointer references. See Memory Management Issues.

18.6. File Variables#

Normally, any open files are closed when the front end exits, but when MAKE_FRONT_END_CALLABLE is used, any open files must be closed if the front-end were to terminate abnormally. This is accomplished by having a special cleanup routine (e.g., templates_cleanup) in any file that opens and closes files. The cleanup routine is present only when MAKE_FRONT_END_CALLABLE is used. File variables must be set to NULL when they do not refer to an open file. The cleanup routine generally calls close_file_if_open, to close any files that may be open. The cleanup routines are called by fe_cleanup (in fe_wrapup.c).

18.7. Command-line Parsing#

cmd_line.c contains code to parse the command line; cmd_line.h contains the associated declarations. The main routine is proc_command_line. Aside from setting switches, proc_command_line establishes the primary source file, the preprocessing output file if there is one, the raw listing output file if there is one, the cross-reference output file if there is one, and the include file search list.

If multiple source files can be compiled, get_next_source_file is called after each source file is compiled to see if there is another file name on the command line. It also does some minor initialization.

proc_command_line is called only once, even if multiple source files are specified on the command line. Since the IL memory for each source file is separate and self-contained, proc_command_line cannot allocate anything in the IL memory: For one thing, it does not exist yet, and for another, there will be a distinct copy for each source file compiled. Therefore, file names and other information saved by proc_command_line must be saved in general memory, and copied into IL memory later (e.g., in fe_init) for each source file.

18.8. Host-dependent Routines#

host_envir.c contains host-dependent routines; host_envir.h contains the associated declarations.

The following routines are used to maintain the list of directory names to be searched for include files:

  • add_default_include_search_path
  • add_to_include_search_path
  • add_to_front_of_include_search_path
  • push_primary_include_search_dir
  • pop_primary_include_search_dir
  • change_primary_include_search_dir

The following routines are used to tear apart and build up file names:

  • directory_of
  • gs_directory_of
  • derived_name
  • replace_file_name_suffix
  • combine_dir_and_file_name

The following routines are used to open files in various ways:

  • open_source_file
  • reopen_source_file
  • okay_as_output_file
  • open_output_file
  • reopen_error_output_file

delete_file is used to delete files (specifically, files that were partially generated and are to be canceled because of detection of an error).

open_temp_file and close_temp_file provide support for creating temporary files. (Note: At the present time, the front end only uses temporary files when generating an intermediate language file for immediate use by a back end called in the same program, and in the C-generating back end.)

normal_termination and term_compilation are used for termination of the compilation. The low-level functions of term_compilation are available individually through write_signoff and exit_compilation.

identify_source_file is called from fe_init at the start of compilation of each source file when there are multiple source files on the command line. It outputs a message identifying the start of the compilation of the file.

set_signal_handlers is called during initialization to set up handling for signals that should abort the compilation.

This file is likely to be modified for different hosts. See the section on host configuration for more information.

18.9. Intermediate Language Traversal#

il_walk.c contains the routines that walk the intermediate language tree, and il_walk.h contains the associated declarations. walk_entry.h contains the detailed code to process the fields of each IL entry kind.

“Walking” the tree means traversing the in-memory form of the IL and calling a supplied routine for each node in the tree. This is used in displaying the IL in human-readable form, in writing the IL to a file, and in reading the IL from a file.

The walking is done on a memory region basis, i.e., for the file-scope memory region or the memory region for a function. The routines for those cases are walk_file_scope_il and walk_routine_scope_il. They call walk_entry_and_subtree or walk_string_entry for each entry, and those in turn call the user-supplied entry processing functions. There are two separate entry processing functions supplied for each traversal: one for strings, and one for non-strings. The main reason for the difference is that string entries do not have fixed sizes, so the string routine has an extra parameter for the size.

Since the tree is actually a graph and has cycles, there needs to be a way to tell when the traversal returns to an entry that has already been visited. This is provided by the il_walk_flag in the prefix that precedes each IL entry. This flag starts out as 0 in all entries. On the first traversal, it is changed to 1 in all entries, and it keeps alternating between 0 and 1 on each subsequent traversal. Therefore, on a given traversal, the entries already seen have one value, and those not yet seen have the other value.

While the tree is traversed, one can also request that all the pointers in the tree be remapped in some way. For example, when writing to an IL file in the alternate form, each pointer is replaced by an entry number. Such remapping can be requested by providing a pointer remap function for the IL walk, or explicitly by calling remap_pointers_in_il_entry or remap_il_header_pointers.

Special handling is required for “orphan” entries. These are entries allocated in the file scope memory region that are referenced only from function scope memory regions. When the parents of such entries are written out and then removed from memory, the entries are orphaned because they are not attached to the rest of the file-scope IL tree.

The orphan mechanism makes lists of such entries so that they can be found during traversal of the file-scope IL.

The array orphaned_file_scope_il_entries heads lists of orphaned entries linked together by a hidden orphan-list pointer that precedes the storage for each file-scope IL entry. (It also precedes the an_il_entry_prefix structure.)

The orphan lists are built up while the function scope memory regions are walked. Each time an entry in a function scope memory region contains a pointer into the file scope memory region, the entry pointed to is recorded as a potential orphan. Later, at the end of the walk of the file scope IL, the entries on the orphan lists are visited. If the entries are not in fact orphans, they will have been previously visited, and the IL walk flag will indicate that fact.

A separate mechanism exists for orphaned lists, specifically the lists of local types and local static variables in function and block scopes. Those are unusual in that the list is completely in the file scope memory region but the pointer to the list is in a function scope memory region entry. Each of the entries on those lists is placed on the orphaned entries list, but there needs in addition to be a way to remember the head-of-list pointers so that the next pointers in the entries can be remapped. That’s done with a list of entries of type a_scope_orphaned_list_header pointed to by the scope_orphaned_list_headers field of il_header. Each entry on the list points to a list of types and a list of static variables preserved from a function or block scope.

Obviously, the code in walk_entry.h must be kept up-to-date whenever the IL definition in il_def.h is changed.

il_walk.c also contains some special-purpose walking routines: walk_declarative_entities_in_scope can be used to walk entities like variables in a given scope; traverse_expr can be used to walk an expression tree; and traverse_statement can be used to walk a statement tree. In each case, user-provided callback routines can be called for each node of the tree.

18.10. Needed Flags#

When the configuration switch MAINTAIN_NEEDED_FLAGS is TRUE, the front end will maintain the needed flag in the source-correspondence entry, and the definition_needed flag in class types and routines. The code to do this is in il_walk.c, walk_entry.h, il.c, and scope_stk.c. The needed flags are a more sophisticated version of the referenced flag. They are set by beginning with all externally-defined variables, static data members, and functions in the compilation unit, and marking those as needed. Then the entities referenced by those are marked as needed, and so on. The resulting needed flag setting indicates whether something is “really” needed, as opposed to referenced from something that is itself unneeded. For example:

static int i;         // i is referenced, not needed
static void f() {     // f is not referenced, not needed
  i = 1;
}
int main () {}        // main is needed, not referenced

The definition_needed flag indicates whether the definition of a class is needed, as opposed to merely a declaration. For example:

struct A { };
struct B { };
A *p;      // forces needed on A, but not definition_needed
B b;       // forces needed and definition_needed on B

There is a similar flag for functions as well.

If DEFAULT_REMOVE_UNNEEDED_ENTITIES is TRUE, and that setting is not overridden via a command-line option, unneeded entities are removed from the IL. Classes that have needed TRUE and definition_needed FALSE are modified to eliminate their definitions. Function-local entities are never affected by this process (they are not removed, and local classes are never turned into declarations). The IL tree after alteration is self-consistent, in that a back end that does not use the “needed” flags will find all the IL entries it expects. A back end that does use the “needed” flags should be sure to use both of them, i.e., definition_needed as well as needed.

The “needed” flags processing and removal of unneeded entries works with and without IL lowering, and for both C and C++. When IL lowering is used, the “needed” flag information is determined on the basis of the lowered code. When the C++-generating back end is used, and the source program has templates, elimination of unneeded entities is disabled, because the template bodies may contain references to apparently unneeded entities. This can be changed, carefully, if one has a situation where all template instantiation will be done by the EDG front end rather than by some compiler operating on the source code generated by the C++-generating back end.

The setting of the “needed” flags is done as follows:

  • At the end of scanning a function definition, if the function is externally defined, pop_scope calls mark_as_needed on the routine, which in turn calls set_routine_definition_needed. These use a specialized version of the walk_entry.h IL-walking code to sweep through the IL tree and set the needed flag on all transitively-referenced entities. If the type of reference requires a complete type, and the type is a class, set_class_definition_needed is called to set the definition_needed flag on the class. If a reference to a routine requires that the routine have a definition, set_routine_definition_needed is called for the routine.

    When a routine’s definition is needed, the definition is also walked, and then the memory region for the function can be disposed of. (mark_to_keep_in_il is also called for the function memory region; more on this below.) That means the memory region can be written to a file and then freed, if an IL file is being used. Conversely, as long as a function is not known to be needed, its body must be retained. If the function is still marked as unneeded at the end of the compilation, it is then clear that the function truly is unneeded. It is deleted and its memory region is freed (without writing it to the IL file, if there is one).
  • When the file scope is being popped from the scope stack, several things are done. First, set_needed_flags_at_end_of_file_scope is called. It visits all variables (including static data members), and calls mark_as_needed for those that are externally-defined. This causes the variables and everything they reference to be marked as “needed.”
  • Then, mark_to_keep_in_il is called for the file scope. This uses another variant of the walk_entry.h IL-walking code to sweep through the IL tree and set the keep_in_il flags in the IL entry prefix and the keep_definition_in_il flag in classes and routines. These flags have meanings similar to the needed and definition_needed flags, but they are set in more cases, to deal with the fact that some entities in the IL tree must be retained (for consistency) even if they are “unneeded.” In this sweep, entities that can and will be removed (based on the needed and definition_needed flags and the context where the entities appear) will not be visited. The keep_in_il sweep also notes whether complete class types are needed, and calls set_class_keep_definition_in_il when appropriate, and whether routine definitions are needed, and calls set_routine_keep_definition_in_il.

    All decisions about removing IL entries from the IL tree are based on the keep_in_il and keep_definition_in_il flags, rather than on needed and definition_needed.
  • Then, eliminate_bodies_of_unneeded_functions is called. It eliminate the bodies (but not the declarations) of functions that are not needed.
  • Then, eliminate_unneeded_il_entries is called. It removes types, variables, and routines in the file scope that are unneeded, and reduces classes to declarations where appropriate (by calling turn_class_definition_into_declaration). Auxiliary parts of the IL data structure, such as source sequence entries and the hidden-name table, are updated as necessary.

The IL-walk process used to set these “needed” flags stops when it encounters an entry with the flag already set. This is desirable for compilation speed reasons, and necessary to avoid loops for parts of the IL structure that are graphs instead of trees, but it also makes certain kinds of problems possible. Variables, routines, and types that are not function-local have the potential to be redeclared after the “needed” flag is initially set. For example, a class A might be declared, then marked as needed, then defined. Or, a function might be defined, then declared to add an additional default argument. In all such cases, The needed or keep_in_il flag would be set already on the entry, and therefore an attempt to set it again would result in an immediate return, rather than a sweep through the updated subtree. To avoid this problem, the subtrees of these redeclarable entities are not swept ordinarily; the appropriate flag is set on the entry itself, but its subtree is not visited. At the end of the file scope, when no more redeclarations are possible, the global flag end_of_file_scope_needed_flags_phase is set to TRUE, and that enables the subtree walks. set_needed_flags_at_end_of_file_scope and mark_to_keep_in_il, as they are doing the needed and keep_in_il walks respectively, look for these delayed-subtree-processing entries, and handle them by clearing the needed or keep_in_il flag and then calling the subroutine to set them again, which this time will visit the subtree. For the needed flag walk, remark_as_needed provides a convenient way of doing this.

When ONE_INSTANTIATION_PER_OBJECT is TRUE, a more complicated set of “needed” flags is maintained. In that mode, each instantiation is put out as a separate object file. There is still only one IL tree, but it is marked so that a back end can sweep through it several times and put out a different “slice” each time, each slice containing one instantiated entity (function, variable, or static data member) plus exactly the set of other entities needed by that entity. This is done via a set of per-instantiation “needed” flags, represented as a bit vector implemented as a linked list of entries of type a_per_instantiation_needed_flags_entry attached to the per_instantiation_needed_flags field of the source correspondence. Each slice is assigned one bit in that bit vector, and if the bit is 1 the associated entity is needed in the slice.

The per-instantiation “needed” flags are tested and set by the macros needed_flag_is_set and set_needed_flag. The global variable needed_flag_bit_number indicates the bit number to be tested or set. If it is zero, the needed flag is tested or set instead of the bit vector.

Each instantiation is actually assigned two bits in the bit vector, the second one used for the per-instantiation versions of the class and function definition_needed flags. These are tested and set by class_definition_needed_flag_is_set, set_class_definition_needed_flag, routine_definition_needed_flag_is_set, and set_routine_definition_needed_flag. As with the needed-flag macros, these macros depend on the value of needed_flag_bit_number and use the definition_needed flags if that variable is zero.

The bit number assigned to a given instantiation entity is recorded in the instantiation_needed_bit_number field in the variable or routine. When mark_as_needed is called for the entity, after the standard “needed” flag processing has been done, another sweep is done to set the proper bit in the per-instantiation “needed” flag bit vectors of all entities referenced from the instantiation. The class and function per-instantiation definition-needed bits are also set in this sweep.

All externally-defined entities that are not instantiations are assigned to the slice numbered 1 (though their instantiation_needed_bit_number fields are left as zero), and a similar sweep is done for them. By the end of the compilation, therefore, every externally-defined entity has been assigned to a slice, and the per-instantiation bits show what must be included in each slice. The “1” slice includes everything that isn’t in an instantiation slice.

Some entities are present in multiple slices. Constants, types, and class definitions appear in each slice as needed. Inline functions are also present in a slice if they are referenced in that slice. In general, however, functions and variables can only appear in one slice, or more precisely, their definitions can only appear in one slice. If other slices reference them, they appear only as declarations in those slices. This causes a problem for static functions and variables, which by definition cannot be referenced from other object files. This is handled by making those entities external with generated names. The sweeping routines will set the static_used_by_instantiation flag in static entities that are used by an instantiation, and IL lowering will turn such routines and variables into external entities.

The simple needed and definition_needed flags end up being in effect the unions of the corresponding per-instantiation sets. That is, if any per-instantiation “needed” flag is set on an entity, then the needed flag will be set, and likewise for the definition-needed flags. Therefore, when the normal processing to remove unneeded entities is done, it will remove only entities that are unneeded in all slices.

18.11. Intermediate Language Display#

il_display.c contains routines to display the IL in human-readable form. il_display.h contains the associated declarations. The code is currently set up to be compiled as a standalone utility program: It reads the IL from a file (by calling the IL read routines), then traverses the in-memory form using the IL walk routines, producing output on stdout. il_display could be set up as a subroutine to be called from within the front end.

The processing is very simple: Using the IL walk routines, it walks the file scope and displays it, then walks each routine scope and displays that. The IL walk routines are told to call disp_entry for each non-string entry, and to call nothing for string entries (string entries are displayed when pointers pointing to them are displayed). Most of the display routines are very straightforward. When pointers are displayed, an attempt is made to show what is pointed to. That is, if the entity pointed to is named, or if it is a string, a constant, or a type, a brief summary of the entity is given so one does not have to look at, say, a type entry to find out that it is the entry for int.

If there are errors in the IL, they are highlighted (e.g., “BAD-TYPE-SPECIFIER-KIND”) and the display continues as well as it can. An internal error is not generated.

Obviously, the code in il_display.c must be kept up-to-date whenever the IL definition in il_def.h is changed.

18.12. Writing the Intermediate Language to a File#

il_write.c contains routines that write the intermediate language to a file, and il_write.h contains associated declarations.

IL writing takes the in-memory IL and writes it to a file for later reading.

start_il_file is called to open the IL file, and finish_il_file to close it. Alternatively, cancel_il_file can be called to delete the IL file and suppress further IL file generation if an error is detected during the compilation.

write_memory_region is called to write out each memory region, i.e., the file-scope memory region or a memory region for a function. It writes the memory region at the end of the intermediate language file.

There are two IL file formats. The switch ALTERNATE_IL_FILE_FORMAT in host_envir.h indicates which is to be used. The file header is the same in both formats:

  • A special string identifying the file as an IL file, and including the version number of the IL (see il_file.h).
  • The number of memory regions.
  • The number of function definitions.
  • The file position of the file index (at the end of the file).
  • The file position of the file-scope memory region.
  • The il_header struct, which is the root of the IL tree.
  • The array of orphan-list pointers.

In terms of processing, the space for the header information is reserved at start_il_file time, but the full information is not known until all regions have been written. At finish_il_file time, the file pointer is repositioned and the final information is overwritten on the reserved space.

The file index is also the same in both formats: It is an array of file positions, such that file-index[i] gives the file position of the start of the information for memory region i. The index is the last thing written in the file. No entry 0 is written; the first entry is entry 1 (i.e., for the file scope), which is at the position indicated by the file index position in the header. The other entries follow. There are as many entries as the “number of memory regions” given in the header.

Between the header and the file index is a sequence of one or more memory regions. Since the file-scope region is not complete until the end of the compilation, the file contains all the function regions first, in order of definition, and then the file-scope region at the end.

In both formats, the information for a memory region begins with the memory region number, and the overall sequence of regions is followed by a zero which serves as an end marker if the regions are read sequentially.

Beyond that, the information for a memory region is quite different in the two file formats.

The default file format involves writing out the memory blocks for each region exactly as they appear in memory – i.e., large blocks, in memory image form. On the reading side, these blocks are read back in, and since they are probably not read in at the same address they had when written out, the IL tree is walked to find and remap all the pointers. A tree walk is also required on the writing side, at least for function scope memory regions, to get the data structure for orphaned entries built. The process is fairly fast on both the writing side and the reading side; although the tree walk takes time, at least the I/O for reading the blocks can be done efficiently, in large blocks.

The information for a region in the default format is as follows:

  • The memory region number.
  • The original memory address for the first block header.
  • The original memory address for the primary scope entry for the region.
  • The total size in bytes of the information following (i.e., the total size of all blocks and their associated headers).
  • For each memory block, the block header and block data. The header contains pointers that point to the next block and to the data (all as original memory addresses).

In the alternate file format, each entry is assigned an entry number (which is recorded in the prefix preceding the entry), all pointers are changed to the corresponding entry numbers, and each entry is written individually to the file, preceded by an indication of the entry kind and by the entry number. On the reading side, the entries are read individually and the pointers in them are changed to the proper new pointer values. This format is less efficient than the default format, because the actual I/O is done one entry at a time for both writing and reading (there is of course buffering under this, but it is still less efficient to make thousands of calls to fwrite and fread than to make dozens of calls). The advantage of the alternate file format is that it allows resizing or other alteration of the entries on the reading side. If one wants to add back-end-specific information, for example, one can do so.

The information for a region in the alternate file format is as follows:

  • The memory region number.
  • An array giving, for all the IL entry kinds (see il_walk.h), the number of entries of each kind in this memory region.
  • A sequence of entries, giving for each entry the entry kind, the entry number, the entry length (only for string entries), the orphan-list pointer (only for non-string file-scope memory region entries), and the entry itself. After the last entry, a zero entry kind marks the end of the sequence.

Entry numbers are assigned from 1 for each kind of table; there can be, for example, a type entry 27 and a constant entry 27. 0 is used to represent a NULL pointer. The entry numbers in each function scope start again from 1, but in addition they have the top bit on, i.e., they effectively start with LONG_MAX+2. Thus there can be identical entry numbers in two function scopes (and that is not a problem), but a file-scope entry number can always be distinguished from a function-scope number.

For string entries, the “entry number” used is an offset into a conceptual block formed by concatenating all the strings of that kind, including the terminating null if any, and with space reserved between the strings for the entry prefix and for any required alignment. The offsets start from 1 to preserve 0 as a NULL pointer. In the array of entry counts, the values for string entries indicate the total size of the conceptual block, in bytes, rather than an entry count. Thus, if there were three strings “abc”, “de”, and “fghi” in a region, the prefix takes 4 bytes, and the host machine requires 4-byte alignment, the strings’ entry numbers would be 5, 13, and 21, respectively, and the entry in the count array would be 25.

18.13. Reading the Intermediate Language from a File#

il_read.c contains routines that read the intermediate language from a file, and il_read.h contains associated declarations.

File reading takes the IL in file form and reads it in to re-create the in-memory form.

The routine il_read reads the file header and the file index, then calls read_memory_region to read the file-scope memory region. Other (function-scope) regions must be read explicitly by calling read_memory_region directly. This allows the back end to read each function, process it, and then free its IL by calling free_memory_region, to reduce overall memory requirements.

With the default file format, the information for each memory region is read with one large read, and the pointers in the region are then remapped to the proper values by calling walk_file_scope_il or walk_routine_scope_il with a pointer remap function. If all the memory blocks are read back in at their original addresses (an attempt is made to do this), the IL walk is not needed.

With the alternate file format, the array of entry counts is read at the beginning of processing the region. For each entry kind, an area is allocated that is large enough to contain all the entries of that kind (its size is the number of entries times the size of each entry). This area is then considered an array of entries of that kind, and the entry number is an index into the area. (For string entries, the “entry count” is the actual total size, and the “entry number” is still an index – with the area treated as an array of characters.) Each entry is then read, placed at the right place in the allocated space, and its pointers changed from entry numbers back to real pointers. Note that no overall tree walk is done, and no special orphan processing is required: Since each entry is read individually, one can do all the processing for each entry when it is read, and be sure no entries were missed.

As each region is read, the associated entry of region_scope_entry in il_header is set to point to the primary scope entry for the region.

18.14. Precompiled Headers#

The precompiled headers (“PCH”) facility is intended to speed up compilations that include relatively stable header files. It does this by writing a precompiled header file that is a snapshot of the front end’s internal state at a point in the compilation process immediately following the initial sequence of #includes. On a recompilation of that file (or another that begins with the same sequence of #includes), the front end checks to see if all the header files are unchanged since the precompiled header file was generated, and if so reads in the precompiled header file to re-establish the front end state as it would be after reading and compiling the header files. Reading in the PCH file takes less time than compiling the header files from source, so there is a net savings in compilation time.

Since the point of this feature is saving time, it is important that the process of writing and reading the PCH files be as fast as possible. One would like to see savings of 80% or the like (comparing the time to restore a PCH file to the time to compile the original source code); an improvement of only 30% or so would barely be worth the extra complexity. In addition, one should not have to pay a large price to write out the precompiled header file; it’s okay to slow compilations a little to get the benefits of precompiled headers on later compilations, but if it takes too long to write the file, the net gain is marginal or negative.

These speed issues suggest some basic design decisions:

  • The information written to the precompiled header file must not require a tree walk, because visiting every entry takes a lot of time. (It also requires code or data structures that describe how to do the tree walk, which exist for IL entries but not for the front-end-only data structures.) Therefore, the information must be written out and read back in as large chunks (rather than one entry at a time), and it must not be necessary to remap all pointers in the information to new values (i.e., it must be possible to bring the information back in at the same address). These requirements also make it possible to exploit memory mapping on systems that support it.
  • Precompiled header files should contain snapshots of compilation states, and not information about individual header files. While it’s tempting to imagine writing precompiled header files for individual header files, the work required in carving out the right information and the description of dependencies so that the information can be written to a file, and the work to bring it back in and merge it with the front end state, will make substantial time savings unlikely.

An unfortunate consequence of this latter design decision is that users are probably forced to change their source code to get any meaningful benefit from the PCH facility. In average source code, the order of inclusion of header files is essentially random. The net effect of the various orders of inclusion is often the same, because of the way header files tend to be written, but it is difficult to deduce that fact, and time-consuming. So we have a dilemma: If we strive for maximum save/restore speed, we will find few “natural” common sequences, and therefore generate many PCH files for little benefit. On the other hand, if we try for maximum ability to reuse precompiled headers from unaltered source files, we lose something significant on save/restore speed.

We’ve chosen to go in the direction of maximum save/restore speed. We believe that users are willing to make small changes in their source code to get large reductions in compilation time, and that given a choice between

  1. You don’t have to change your source, and precompiled headers will give you a 15% savings, or
  2. You may have to reorder your initial #includes or add the equivalent of

    #include <everything.h>

    to the beginning of your source files, and precompiled headers will give you a 50% savings,

users will choose option 2. This, of course, is a matter of opinion, and the percentages are fictitious. But for better or worse, this is what we decided.

Because PCH processing is not likely to be helpful on arbitrary files, it is disabled by default.

The code for precompiled header processing is mostly in pch.c, with associated declarations in pch.h.

18.14.1. Prefix and Header Stop Point#

The initial sequence of the source code of the primary source file that is subject to the PCH optimization is referred to as the prefix. It consists of the preprocessing directives preceding the header stop point. The header stop point is determined as follows:

  • Start at the first token of the primary source file that is not part of a preprocessing directive (or the end of the file if there is no such token).
  • If there is a #pragma hdrstop preceding that token, back up to the pragma position.
  • If the position determined by the first two steps is inside a preprocessing directive like an #if, back up to (just after) the previous top-level preprocessing directive.

For example:

/* Comments. */
#define flag 1
#if flag
#include "abc.h"
#else
#include "def.h"
#endif
#include <stdio.h>
/* header-stop point. */
int i = 1;
#include "abc.h"
/* header-stop point. */
#if xyz
#include "def.h"
#pragma hdrstop
#endif

Note that, while determining the prefix, #ifs and the like are not evaluated, so code that will be skipped in the actual compilation can nevertheless serve as a header stop point. This is particularly important with #pragma hdrstop.

At the beginning of compilation, if precompiled headers are enabled, the prefix for the primary source file is accumulated. This is done by opening the primary source file, reading the beginning of it in a special mode while building an in-memory representation of the preprocessing directives read, and closing the source file. This is a separate process from the normal compilation of the file, and is a necessary first step whether the front end ends up using an existing PCH file or creating a new one. See build_prefix_information.

The prefix scan is done by setting building_pch_prefix and calling get_token once. That causes the scanning of all of the preprocessing directives preceding the first “real” token of the primary source file. The preprocessing directives are not evaluated in the usual way; they are just placed on the prefix event list. (In particular, #include directives are just accumulated like the others; the included files are not read.) Each such event contains an enumeration code for the preprocessing directive and a null-terminated string giving the text of the directive following the directive keyword.

After the initial prefix list is built up, the list is trimmed down by discarding anything after a #pragma hdrstop and then, if the end is inside an #if or the like, discarding preprocessing directives back to the last top-level directive.

If the primary source for the compilation is coming from stdin instead of a file, the prefix construction cannot be done (and therefore PCH files cannot be created or used).

A list of the command-line options for the current compilation is also built. It is also a linked list, and each entry on the list indicates one command-line option. Options that have no effect on the state preserved by PCH processing are not put on the list, but the rest (most of them) are. The option code is represented by an enumeration, so the difference between the single-letter form of the option and the keyword form is erased, but in all other respects the list built up will have to match the corresponding list for a PCH file exactly – same options, same order.

18.14.2. Using a PCH file#

Once the prefix has been created, the front end will look to see if any of the existing PCH files can be used to avoid some part of the compilation for the current file. A host-dependent routine is used to fetch the names of all PCH files in the appropriate directory (the directory can be specified via the --pch_dir command-line option). Each file is checked to make sure it is a regular file (not, say, a directory), that it has the right header string identifying it as a PCH file, that the compiler version number and generation date/time are the same, and that the flag indicating that the file is completely written has been set. If any of those checks fails, the file is closed and not considered further; in effect, it’s considered not to be a PCH file.

The next round of checks seeks to determine whether the PCH file is applicable to the current compilation. The front end checks that the current directory and the one recorded in the PCH file are the same; that the command-line event lists are the same; and that the preprocessing directive event lists are the same (or the current file’s prefix list is longer than the prefix list in the PCH file). If those are okay, the front end checks that the header files included in the PCH file have not changed since the PCH file was built (this is done by checking the timestamps of the files against the timestamps recorded in the PCH file). If the timestamps indicate that the PCH file is out of date, it is deleted.

A PCH file that makes it this far can be used, but it might not be the best choice: There might be another PCH file that also matches but includes more of the prefix preprocessing directives. Therefore, if the match is not perfect, other PCH files are considered and the best-matching PCH file is chosen. Each PCH file is closed after it is considered and must be reopened later if selected as the best file.

If the user has used the --use_pch command-line option, the indicated PCH file is chosen directly. No other PCH files are considered, and the applicability test is not done (at this point; see below).

Once the best-matching PCH file is selected for a file called (say) x.C, if the PCH file selected is not x.pch, and there exists an x.pch, x.pch is deleted. This is a simple way of getting rid of PCH files that are no longer needed.

Once a PCH file is selected, the front end state is restored from the file. The PCH file is opened again and the prefix check is done again to make sure the file has not changed since the first check. For a PCH file specified via a --use_pch command-line option, this is the only applicability test done (and if the file fails, a warning is issued and the compilation continues without using the PCH file); for other cases, this is a repetition of the applicability test already done. Next, the global variables registered via register_pch_saved_variables are read and restored. Finally, the contents of the various memory regions (IL and front end) are restored. These are restored at the original addresses, so it is not necessary to walk through the restored data to adjust pointers. (See the section on memory management issues below for more information on this aspect.)

The restoration process occurs after most initialization of the front end, so some variables are initialized then overwritten by the restoration from the PCH file.

Once the front end state has been restored from the PCH file, the lexical routines are reset and the primary source file is opened again. The prefix is skipped by doing a get_token in a special lexical mode. This is very similar to the scan done to accumulate the prefix, in that initial preprocessing directives are passed over uninterpreted. However, once the preprocessing directive that marks the end of the matched part of the prefix is processed, normal processing of preprocessing directives is re-enabled. Shortly thereafter, the initial get_token finishes, and normal compilation continues from that point (we know we are between top-level declarations, so we can continue simply by entering the main compilation-unit loop).

The PCH restoration process restores the front end state after compilation of the precompiled headers, but it doesn’t duplicate the output that would have been done on auxiliary files during the skipped part of the compilation. Diagnostics that would have been written, and output lines on the raw listing and cross-reference output files that would have been generated, are not put out when a PCH file is used. If the front end is configured to use an IL file, that file also is not generated as this point. However, the IL file does get generated by a catch-up process done after the point where a PCH file is generated or a decision is made not to generate one; see below.

Because the source sequence numbers of the original compilation may not exactly match the sequence numbers in the current compilation (e.g., because of whitespace differences), the source file mapping information in the IL in the PCH file is restored as, in effect, the first include file on the list under the primary source file of the current compilation.

18.14.3. Creating a PCH File#

After the front end has read in a PCH file or chosen not to use one, the “normal” compilation process begins. When the compilation reaches the header stop point, we have the option of creating a PCH file that encapsulates the front end’s state at that point for later restoration.

Note that it is possible to both use and create a PCH file in the same compilation, because the best-matching PCH file might contain only the head of the full prefix. In that case, the internal state is restored from the PCH file, then the rest of the prefix is compiled in the normal way, and at the header stop point the compiled form of the full prefix is saved to a new PCH file.

The position of the header stop point was determined by the prefix scan at the very beginning of the compilation. Its position is defined in terms of a preprocessing directive, more specifically as the end of the processing for that directive. Compilation starts up or continues in a mostly normal way, with the header stop position set as a sort of breakpoint. When the end of the indicated preprocessing directive is reached, the front end will check to see whether a PCH file can be generated.

On the way to the header stop point, some things may happen to preclude generation of a PCH file:

  • A reference to one of the macros __DATE__ or __TIME__. Since these would change from compilation to compilation, there is no point to saving the state of a compilation that uses them.
  • A #pragma no_pch. This is an explicit request not to generate a PCH file.
  • A #line directive. (Such directives could probably be allowed but are not accommodated at present.)

When the header stop point is reached, generate_precompiled_header is called. It starts by making several checks that the current position is acceptable:

  • The current position is not inside a declaration, i.e., it is between two file-scope declarations, or before the first such declaration, or after the last. next_token_is_top_level_decl_start, maintained by the declaration-scanning routines, gives this information.
  • The preprocessing #if stack is empty, i.e., the current position is not inside a conditionally-compiled section.
  • There have been no errors.
  • The global variable cannot_create_pch_file is not set. It is set for things that preclude generation of a PCH file (e.g., referencing __DATE__).

If the current position is acceptable, a precompiled header file can be written, and the front end will write out a file xxx.pch, where “xxx” is the base file name of the file being compiled. The user can specify a different file name via the --create_pch command-line option. If a file with the chosen name already exists, it is deleted. Precompiled header files are written in the current directory by default, but another directory can be specified via the --pch_dir command-line option.

The following information is written to the PCH file, in the order given:

  • An string that identifies the file as a PCH file, and gives the front end version number and generation date/time.
  • The current directory name (this may affect the include file search order, etc.).
  • The list of command-line events, i.e., all the command-line options that could have an effect on PCH files.
  • The list of preprocessing directive events, i.e., the sequence of preprocessing directives whose compiled form is encapsulated in the PCH file.
  • Timestamps for all the header files referenced by the compilation (this includes all files, not just those included directly from the primary source file).
  • The values of all front end variables registered with register_pch_saved_variables.
  • The contents of all memory regions, i.e., the front end memory region and all IL memory regions.
  • Once the rest of the information has been written, the front end goes back to the header and writes a flag indicating that the PCH file is complete. This flag prevents other compilations from using a PCH file that is not yet completely written.

After writing out the PCH file, the front end continues with compilation. Because the header-stop point has been passed, no further precompiled-header checking or processing is done in the rest of the compilation.

If an IL file is being generated, function scope memory regions that are completed before the header stop point are held in memory and neither written to the IL file nor freed. Once the header stop point has been reached, and a PCH file has been written or a decision made not to write one, those memory regions are written to the IL file and freed.

18.14.4. Memory Management Issues#

As was mentioned earlier, one of the prerequisites for fast PCH processing is the ability to bring memory regions back into memory at the same addresses they had when written out. Clearly, that requires that (a) nothing else happens to have been allocated at the desired addresses, and (b) there is some way to request allocation or assignment of those addresses so that no later allocation uses them.

The front end provides two schemes that accomplish this goal. The first and better scheme requires memory mapping support from the underlying operating system (UNIX and Windows NT versions are provided). A fixed set of addresses is chosen for all memory regions (the exact addresses are either left to the discretion of the operating system or chosen in the configuration of the front end; see USE_FIXED_ADDRESS_FOR_MMAP). Initially, these memory addresses are mapped onto a temporary file. When a PCH file is created, the memory is written to the PCH file (using normal file I/O). On restoration of a PCH file, the default mapping to the temporary file is released, and the reserved memory addresses are mapped onto the PCH file being restored. This brings the saved information into the front end’s address space without any immediate I/O. (In the actual implementation, the memory addresses are tentatively mapped once to verify that the mapping can be done, and then, once it is known that all the mappings will be successful, they are done again with the official mappings.) All dynamically allocated memory that is not part of a memory region is just allocated in the normal way through malloc, since it need not be at any particular addresses. This means, in effect, that there are two completely separate pools of memory – memory region allocations are done at the fixed mappable addresses, and everything else is done more casually in the traditional way.

The other implementation does not require memory mapping, but it is slightly less friendly to users. Very early in the execution of the front end a large amount of memory is preallocated in large blocks and saved off to the side for use exclusively for memory regions. This is done early so that the allocated addresses are reproducible (i.e., one can get the same addresses on all compilations). If the allocations were done later, their addresses would be affected by minor variations in the amount of memory used during the command-line processing, and perhaps by memory allocated by runtime routines of the host C compiler (e.g., printf formatting). As the compilation proceeds, requests for memory for memory region blocks are satisfied by taking blocks from this preallocated list. All other allocations (for things not in memory regions) are done directly through malloc. When a PCH file is created, the memory addresses of blocks are written to the PCH file along with the block contents. Also written is an array of the preallocated memory block addresses used, serving as a sort of allocation history (this array is written to the file also in the memory-mapped mode). Once the header stop point is passed, and either a PCH file was written or a decision was made not to write one, the remaining preallocated memory is freed, and subsequent allocation for all purposes is done directly through malloc.

If the preallocated space is too small to hold the memory regions created up to the header stop point, the precompiled header cannot be created (since the remaining required space cannot be allocated at predictable addresses). A warning is issued, telling the programmer that the --pch_mem option can be used (on a subsequent invocation of the front end) to request a larger preallocation to avoid the problem. The compilation continues normally – the PCH file cannot be created, but the compilation can be completed regardless.

Since the preallocated blocks have a fixed (though large) size, any requirement to allocate a single entity larger than that fixed size cannot be satisfied from the preallocated pool, and therefore precludes generation of a PCH file. This could happen, for example, if a user program contains a single literal string containing 200,000 characters. Again, the compilation continues without difficulty, but a warning is issued that a PCH file could not be generated.

When a PCH file is read back in the non-memory-mapped mode, the allocation history information history is read in from the PCH file and used to check that the preallocated blocks are indeed at the same addresses (on some systems, in spite of the care taken to allocate the blocks early, the addresses still change; for example, under MS-DOS, the addresses can be affected by placing a new TSR program in memory). If the addresses are not right, it is still possible to give up on using the PCH file and continue with normal compilation. If the history information matches, the memory regions are read in: The address for each block is read from the file, the preallocated block at the right address is assigned, and the memory region contents are read from the PCH file into the proper addresses.

It should be noted that if a given environment provides a (nonstandard) way of allocating memory from a pool distinct from the normal malloc allocation pool, the non-memory-mapped approach can be made to work without preallocation and without the --pch_mem command-line option.

When multiple source files are specified in one invocation of the front end, the preallocation scheme breaks down: The first file can be processed okay, but on subsequent files the preallocated memory has been used and reclaimed and is no longer in quite the right predictable state. This can probably be solved in some way, but the current implementation takes the easy way out and disallows PCH file use on compilations where more than one source file is specified on the command line.

18.15. Predefined Macro Definition File#

Macros may be predefined either by explicit calls of enter_predef_macro, and/or through use of a predefined macro definition file. Whether or not the front end reads a predefined macro definition file is controlled by the DEFAULT_USE_PREDEFINED_MACRO_FILE macro. The file EDG_BASE/EDG_AUXILIARY_INFO_DIR_NAME/PREDEFINED_MACRO_FILE_NAME is read to provide the macro definitions. The EDG_BASE directory can be specified by an environment variable (EDG_BASE), build-time macro (DEFAULT_EDG_BASE), or command-line option (--edg_base_dir). EDG_AUXILIARY_INFO_DIR_NAME and PREDEFINED_MACRO_FILE_NAME are configuration macros with default values of “lib” and “predefined_macros.txt”. If EDG_AUXILIARY_INFO_DIR_NAME is null, the EDG_BASE directory is used directly. The format of the entries in the file is:

mode,!mode,mode   cannot_redefine   macro_name   macro_value
  • mode is a label from the predefined macro modes table. The macro is defined if the mode is set, or if the mode is not set when !mode is used. The macro is defined if any of the mode tests is TRUE. The set of mode values is specified in host_envir.h.
  • cannot_redefine indicates whether the predefined macro may later be redefined. The value must be “yes” or “no”.
  • macro_name is the name of the macro to be defined.
  • macro_value is the value to which the macro should be defined. All of the characters until the end of the line are used as the macro value.

For example:

gcc no  __HAVE_BUILTIN_SETJMP__ 1
gpp no  __CHAR_BIT__ 8
gnu no  __GNUC_PATCHLEVEL__ 0
gnu no  __WINT_TYPE__ unsigned int
gnu,microsoft __IS_GNU_OR_MICROSOFT_MODE 1

A script (make_predefined_macro_file) is provided that automatically generates the predefined_macros.txt file for a given version of gcc/g++ or clang. It does this by invoking gcc and g++ or clang and capturing the predefined macro information.