6. Symbol Table#

The primary function of the symbol table is to provide a mechanism for associating names and entities: a name can be entered into the symbol table in association with a particular definition and then later looked up to get back to the same definition. Name visibility is controlled by an implementation of nested name scopes and class member inheritance. Provision is also made for overloaded operators and user-defined conversion functions, which are not looked up by name. In addition, the symbol table serves as a repository of information that is needed for front-end processing but is of no use to back ends. symbol_tbl.c contains the code to manage the symbol table, lookup.c contains the code that looks up identifiers, scope_stk.c contains the code that manages the scope stack, and symbol_ref.c contains the code that manages the recording of references to symbols. symbol_tbl.h, lookup.h, scope_stk.h, and symbol_ref.h contain the associated declarations.

6.1. Basic Organization#

6.1.1. Symbols and Symbol Headers#

The key components of the symbol table are entries of type a_symbol and a_symbol_header. Typically, symbols represent specific uses of a name (e.g., the same name may be declared in different name scopes or belong to different name spaces), but all symbols for a particular name share the same symbol header. Each symbol entry points to a header which identifies the name associated with it; the symbol header points in turn to symbols that share its name. Symbols with the same name from different translation units (see Multiple Translation Units) share the same symbol header.

For certain special cases (overloaded operators and user-defined conversions) the symbol header has no actual name but can be thought of as representing a family of virtual names to be constructed for the associated symbols: a single symbol header is shared by all symbols that overload a given operator or, for user-defined conversions, a given return type. (The names generated by the compiler to assure type-safe linkage are not produced by the front end and so do not appear in the symbol table.)

At the gross level the symbol table has three distinct pieces, each of which has its own “root”, a global variable that is accessed in the lookup process:

  • For named symbols symbol_table is a hash table where each bucket points to a linked list of symbol headers. The name is reduced to a hash value that is used to index the array.
  • For overloaded operator symbols opname_symbol_table is an array of symbol header pointers that is indexed by members of the enumeration an_opname_kind. There will be one symbol header for each overloaded operator.
  • For user-defined conversion symbols conversion_header_list points to a linked list of entities of type a_conversion_header, one for each result type for which the user has declared a conversion routine. Each conversion header points to a single symbol header.

6.1.2. Active, Inactive, and “Other” Symbols#

Each symbol header contains pointers to three lists of symbols of the same name (or that otherwise share the symbol header): the active symbol list, the inactive symbol list, and the other symbol list.

6.1.2.1. The Active Symbol List#

The list of active symbols includes all symbols that are still in scope. The order of the symbols on the list is significant since certain of the lookup mechanisms take the first symbol found. When symbols are entered into the symbol table, their entries are always placed at the beginning of the list, so that declarations in inner scopes will eclipse outer-scope declarations of the same names. An exception is made when a tag and a nontype with the same name exist in the same scope in C++. In such cases the tag is always placed on the list following the nontype so that it will be hidden except when accessed using an elaborated type specifier. When the inner scope ends, its symbol entries are removed, which exposes the outer declarations.

Note, however, that symbols from the same scope can coexist on the active list, since different kinds of symbols can belong to different name spaces. For instance, a variable, a macro, and (in ordinary C) a tag can all share the same name. The array name_space_for_symbol_kind maps a symbol kind into a_name_space_kind (e.g., nsk_label, nsk_macro, nsk_other). The lookup mechanism must check for name-space compatibility: in most cases, it takes the first in the active list whose kind maps into the appropriate name space. The name space kind nsk_other is the normal name space for types, variables, enumeration constants, and functions.

6.1.2.2. The Inactive Symbol Lists#

When a scope terminates, the symbol for each name declared in that scope is removed from the active list of the associated symbol header. In most cases, this means the symbol is removed from the symbol table (i.e., that it can no longer be found by any lookup mechanism). But for some symbols (e.g., class and namespace members) it must remain possible to look up a name even after it passes out of scope, and so its symbol is moved from the symbol header’s active list to its inactive list. The order of symbols on the inactive list is not significant.

Thus, the inactive symbol list will comprise symbols for class members that are no longer in scope but are still indirectly accessible, whenever the class is accessed.

Similarly, when the template declaration scope terminates, its symbols representing the template parameters are also added to the inactive list. This enables them to be accessible during instantiation of the template.

Namespace scopes are not terminated when the namespace is closed because the namespace may be extended later in the compilation. Namespaces may be extended either explicitly by opening a namespace extension scope, or implicitly as a side-effect of a namespace member definition in the file scope or as a side-effect of the instantiation of a template entity defined in a namespace scope. When the original namespace definition scope is closed, the namespace members are moved from the active list to the inactive list. If the namespace scope is later extended, new symbols for the namespace are entered directly on to the inactive list.

The file scope is popped when the scanning of a translation unit is completed so that its symbols are not on the active list if any other translation units must be processed. As with namespaces, the symbols are moved to the inactive list to permit the file scope to be reactivated later, if additional processing of the original translation unit is required (e.g., to generate template instantiations).

6.1.2.3. The “Other” Symbol List#

The “other” symbols list is used for symbols that are never found using the normal lookup routines. Specifically, the “other” symbols list includes “extern” symbols and synthesized namespace projection symbols, both of which are described below.

6.1.3. Scope List#

When symbols are entered into the symbol table, they are also entered onto a list of the symbols declared in the current name scope (discussed in detail below). [1] This list preserves the names in declaration order and is referenced in removing symbols from the active list (and usually from the symbol table) when the name scope is terminated.

Typically, this list lasts no longer than the scope stack entry that points to it – when the entry is popped off the scope stack, the associated list is discarded. In some cases, however, the list must be kept around for later use.

  • For class scopes it is stored in the class symbol supplement for the class for use in cases in which the declaration order must be observed (e.g., in doing member initialization).
  • For function prototype scopes it is used to reenter parameter symbols in the function scope.
  • For namespace scopes it is preserved in the pointers block for the namespace.
  • For file scopes it is preserved in the translation unit entry (there can be more than one file scope when processing multiple translation units).

6.1.4. Special Cases#

There are several classes of symbol for which special treatment is required.

Symbol entries for constructors, though they have names, are not entered directly into the symbol table. That is, they do not appear on the active or inactive list of the symbol header for the name. Rather, they are accessed from the constructor field of the class’s supplement entry. Thus the symbol lookup for constructor X::X is always by way of the symbol supplement for class X.

A tagless class, struct, or union presents an interesting case because it has no name but must have a symbol to serve as a repository of information for the front end’s processing. Such symbols all share a single symbol header (see static variable unnamed_tag_symbol_header in symbol_tbl.c) but are not entered on any list.

C++/CLI property accessor functions are members of a property which is itself a member of a class, but the property is not a nested class; it’s a data member. For the accessors, special symbol headers are generated that are a combination of the symbol header for the property and the symbol header for the access name. See get_property_or_event_accessor_symbol_header. Event accessors are handled similarly.

Error symbols are created in various cases when errors are detected. They have no name associated with them, and, like unnamed class symbols, they all share the same symbol header (see static variable error_symbol_header in symbol_tbl.c).

6.1.5. Symbols and IL Entries#

For the most part symbol entries are little more than an indication of the kind of symbol and a pointer to an IL entry that provides more information. Thus, sk_variable and sk_static_data_member symbols point to an entry of type a_variable; constants, types, tags, fields, functions, and labels are similar. The principal function of most symbol entries is simply to indicate a mapping between names and the intermediate language constructs to which those names correspond.

But symbol entries also serve as repositories of information that the front end requires for its processing but that back ends do not care about. For instance, sk_keyword symbols contain the keyword token value and for sk_macro symbols there is a pointer to a_macro_def. Most notably, symbols representing classes, structs, and unions point not only to the corresponding IL entry but also to an entry of type a_class_symbol_supplement, which contains a number of flags and pointers (including a pointer to the constructor symbol, as described above).

Some symbols just point to other symbols: sk_projection symbols, which represent names in a derived class that are inherited from a base class, contain a pointer to the original symbol, and sk_overloaded_function symbols point to a linked list of routine symbols that represent instances of the overloaded name.

sk_parameter symbols are created during parameter processing. They are needed because parameter names hide names from containing scopes (e.g., in default argument expressions). However, they do not refer to any IL entity; rather, when a function definition is present, they are reentered into the function scope and turned into sk_variable symbols which do point to variable entries.

6.1.6. Function Overloading#

Each instance of an overloaded function is represented by a symbol – of kind sk_member_function for member functions and sk_routine otherwise. However, these symbols are not directly present in the symbol table. Rather, the whole family of functions that share the common name or operator in a given scope is represented by an sk_overloaded_function symbol, and that is what actually appears in the symbol table (i.e., is on the active or inactive list of a symbol header). Consequently, name lookup finds only the overloaded function symbol, not the symbol for the particular function required. [2]

The actual function symbols appear on a linked list pointed to from the overloaded function symbol. The order of symbols on the list is not significant, since the argument matching algorithm requires visiting every instance anyway. The list may contain function symbols, function template symbols, projection or namespace projection symbols that point to functions or function templates, or a combination of any or all of these.

6.2. Name Scopes#

6.2.1. push_scope and pop_scope#

Aside from the structure described above, the symbol table also has a name scope structure. Each time a new name scope is entered (for the file scope, a class, struct, or union, a function, a block, function prototype, namespace, template declaration, or template instantiation), push_scope is called to push a new entry onto the scope stack. (See scope_stack, a global variable that points to an array of elements of type a_scope_stack_entry. Since the scope stack is implemented as a dynamically allocated array that can be extended by reallocation when necessary, elements should always be accessed by indexing into the array rather than by reusing a pointer.)

There are actually a number of different routines used to push different kinds of scopes (e.g., push_template_instantiation_scope). The various routines all call push_scope_full, which does most of the processing. push_scope is used for pushing “normal” scopes that don’t require that extra arguments be supplied when called. This section uses the name push_scope to refer to the set of functions that add scopes to the scope stack. Most of the actions described here are performed by push_scope_full.

push_scope assigns a unique number to each scope, and when an identifier is entered into the symbol table, the scope number is recorded in its symbol (field decl_scope); this number may be referenced during symbol lookup, and it is used to detect multiple declaration errors during symbol entry. When the name scope ends, pop_scope is called to remove its symbols from the symbol table (class member symbols are an exception, as described above) and to pop the scope stack entry off the stack. pop_scope calls wrapup_scope, which removes the symbol entries for the scope from the active list and reenters them on the inactive list when appropriate (for class, namespace, and template declaration scopes). wrapup_scope also calls end_of_scope_symbol_check, which issues diagnostics on incomplete types and on entities that were declared and never referenced or set and never used. When the file scope is popped at the end of the translation unit, pop_scope calls wrapup_namespace_scopes, which in turn calls wrapup_scope for all of the namespace scopes in the translation unit.

The previous_scope field of the scope stack entry is used to determine the order in which scopes on the scope stack are to be considered for name lookup purposes. For most scopes, this field is set by push_scope_full.

push_scope and pop_scope maintain a number of state variables, some global and some local to symbol_tbl.c:

  • depth_scope_stack gives the depth of the current name scope in the scope stack. [3]
  • decl_scope_level gives the scope depth of the scope entry with which declarations should be associated. [4]
  • depth_innermost_function_scope is the depth of a sck_routine scope stack entry that is or contains the current scope; if there is none or if a class scope intervenes its value is NO_SCOPE_DEPTH.
  • num_classes_on_scope_stack is the number of class entries currently on the scope stack (including sck_class_reactivation as well as sck_class_struct_union scopes).
  • depth_of_innermost_scope_that_affects_access_control is used in determining whether a piece of code has member access privileges to a given class (e.g., in a friend function).
  • depth_innermost_namespace_scope is the depth of the innermost namespace scope that is visible on the scope stack, or the depth of the file scope if no namespace scopes are visible.
  • depth_innermost_instantiation_scope is the depth of the current template instantiation scope or NO_SCOPE_DEPTH if no instantiations are active.

push_scope also causes new memory regions to be created for the file scope and for function scopes. When the scope is popped, the following processing is done:

  • If IL lowering is being performed, should_delay_lowering_on_function is called to determine whether the function should be lowered at this point. A number of circumstances can cause lowering to be delayed, such as the need to preserve unlowered IL for constexpr evaluations, the need for a module-id that has not yet been determined, etc.
  • should_delay_finishing_of_function_body is called to determine whether a function body should be “finished” at this point. If so, finish_function_body_processing is called. This will lower the function, if it has not already been lowered, do needed flag processing (if necessary), and any other operations that may be needed before the memory is potentially freed.
  • check_for_done_with_memory_region is called either to write out the entries in the memory region to a file and reclaim the entire region for reuse or, if the interface to the back end is in memory, to reclaim only the unused portion of it. In some cases the memory must be kept and IL file generation deferred. This is done, for example, when a function body is kept for potential inlining, or if the function contains a generic lambda that may be instantiated later.

The name scope for a class whose definition is completed must be reinstated to process the definition of a member function of that class or the initializer of static data member of that class. This is accomplished by calling push_class_reactivation_scope, which has the effect of making the symbols on a symbol header’s inactive list visible (that is, those inactive list symbols whose decl_scope value matches the scope number of the reactivated scope become visible). When the class to be reactivated is a nested class, push_class_reactivation_scope reactivates its enclosing class first to create the required name visibility and hiding. If the class being reactivated is from a namespace, the namespace will also be reactivated (see below). A matching call to pop_class_reactivation_scope is required once processing for the member function definition or static member initializer is finished.

The assoc_type field in sck_class_struct_union and sck_class_reactivation scope stack entries points to the type entry for the corresponding class, struct, or union.

Namespace scopes, like class scopes, are reactivated when processing the declarations of namespace members outside of their namespace. As mentioned earlier, namespace scopes are unique in that they are not actually closed when the end of a namespace definition is reached. The namespace can be extended later either by a namespace extension declaration or by a name that is injected while processing the definition of a member of a namespace outside of the namespace (including the instantiation of templates defined in the namespace). Namespaces are reactivated by calling push_namespace_reactivation_scope. pop_namespace_reactivation_scope is called when the reactivated namespace is no longer required. Reactivating a namespace causes any namespaces that enclose the reactivated namespace also to be reactivated. Namespaces are extended by calling push_namespace_extension_scope. When the namespace extension scope is no longer required, pop_namespace_extension_scope must be called. Extending a namespace causes any namespaces that enclose the extended namespace also to be extended. A namespace is reactivated in contexts in which the names from the namespace must be visible but new names cannot be added to the namespace. A namespace is extended in contexts in which names from the namespace must be visible and new names can be added.

The scope stack contains a set of pointers that are used to track the symbols for the scope and the “last” pointers that point to the end of the each of the lists of IL entries associated with the scope. For most scope kinds, this information is no longer needed when the scope stack entry is popped off of the stack. Because namespace scopes can be extended, this information must be preserved for those scopes. Furthermore, a given namespace can appear on the scope stack more than once when templates are being instantiated. When more than one scope stack entry exists for a given namespace, they must share the same set of symbol and IL entry pointers. The pointers that must be preserved between extensions of the namespace and shared between scope stack entries for the namespace are in a structure called a_scope_pointers_block. Each scope stack entry contains a pointers block structure and field named assoc_pointers_block that points to a pointers block structure. For all scopes except namespace and namespace extension scopes, the assoc_pointers_block pointer is NULL, indicating that the pointers block within the scope stack entry should be used. For namespace and namespace extension scopes, the assoc_pointers_block pointer points to a pointers block in the namespace symbol supplement for the namespace. The macro assoc_pointers_block_of should be used to obtain a pointer to the appropriate pointers block for a given scope stack entry.

push_template_instantiation_scope is called when the appropriate context in which a given template can be instantiated needs to be established. When a template instantiation scope is pushed, the sequence of scopes used for name lookup purposes is changed. In fact, there are four distinct sets of scopes used for name lookup when a template instantiation in progress:

  • the scopes associated with the local context of the template: These include the template instantiation scope (with which the template parameters are associated), the scope associated with the class or function being instantiated, and any scopes nested within the class or function being instantiated.
  • the template definition context scopes: The scopes that are only part of the context in which the template was defined.
  • the referencing context scopes: The scopes that are only part of the context in which the template reference that caused the instantiation occurred, beginning with the innermost namespace scope.
  • the common scopes: The scopes that enclose both the definition context and the referencing context. This always contains at least the file scope.

push_instantiation_context determines which of the required namespace scopes are already present on the scope stack and pushes those entries that are not already present. It then calls reactivate_class_and_instantiation_scopes to reactivate any classes that may enclose the template that is being instantiated. If an enclosing class is itself a template class, an instantiation scope for that class is also pushed so that its template parameters will be visible to the entity being instantiated.

fixup_instantiation_scopes is called after all of the scopes have been pushed. It performs the following functions:

  • All of the instantiation scopes that were pushed, except for the outermost one, are flagged as nested instantiations. Only nonnested instantiation scopes are considered for special instantiation context lookups, as described later.
  • The context and common scope depths are recorded in the nonnested instantiation scope.
  • The previous scope pointers of the outermost definition context scope and the referencing context scope are set to point to the common scope.
  • The next_scope_that_affects_access_control field is updated to make sure that it does not point at a scope that is not part of the current context.
  • In the set of instantiation scopes pushed by a call to push_template_instantiation_scope, all except the innermost one have their exclude_from_context_output flag set. This prevents unnecessary context information from being generated during diagnostic output.

The previous_scope field is also updated when a namespace reactivation scope is pushed within a template declaration scope. In such cases, the names from the namespace extension scope must not be considered until after the template parameters. For example, in the following case the T in the function parameter list must refer to the template parameter and not N::T.

namespace N { typedef int T;}
template <class T> void N::f(T);

If GENERATE_SOURCE_SEQUENCE_LISTS is configured to TRUE, source sequence lists for active scopes are recorded in the scope stack entry (see source_sequence_list and end_of_source_sequence_list in a_scope_stack_entry). When pop_scope is called for the file scope or a function scope, the list is moved to the corresponding IL scope. Otherwise, the list is merged into the list of a containing scope, which usually means that the list is simply appended to the list of the immediately containing scope; for sck_template_instantiation scopes, however (when a compiler-generated specialization represents an automatic instantiation – that is, when either CLASS- or NONCLASS_TEMPLATE_INSTANTIATIONS_IN_SOURCE_SEQUENCE_LISTS is configured to TRUE), insert_instantiation_src_seq_list is called to move the list to an appropriate location in the list of an enclosing scope.

6.2.2. Connection With IL Scope Entries#

While not all scope stack entries correspond to an intermediate language scope entry, all intermediate language scope entries have a corresponding scope stack entry. Every scope stack entry has a (possibly NULL) pointer to an intermediate language scope entry; every intermediate language scope entry refers back to the corresponding scope stack entry by means of a scope number.

For example, a new scope stack entry is pushed for a statement block whether or not it contains any declarations, but the intermediate language scope is not created until it is needed, and not till then is the IL_scope pointer in the scope stack entry updated.

The scope stack entries for the file scope, namespaces, classes, routines, and blocks are closely tied to a corresponding intermediate language scope entry: the latter has pointers to the lists of types, variables, etc. for that scope, and the pointers block has the pointers to the ends of those lists. The pointers block in the scope stack entry is used for all scopes except for namespace and namespace extension scopes, for which the pointers block is located in the namespace symbol supplement. These tail pointers are only needed while the scope is active, to add entries at the end of the list, and having them in the pointers block reduces the size of the intermediate language scope entry.

6.3. Classes#

Classes (in the generic sense, including structs and unions) require the management of quite a lot of information. Some of this must reside in the intermediate language. For instance, the entities that represent members of a class are kept on lists pointed to from the intermediate language scope entry associated with the class; and the base classes from which the class is derived are associated with the class’s type entry. In these and other cases information associated with a class is stored in intermediate language declarative constructs, and appropriately so, since the information is of use to back ends.

However, there is a great deal of information that the front end needs but back ends do not care about, and it would be wasteful to store it in the intermediate language. Instead, it is kept in the symbol table so that it can be thrown away when front-end processing is complete.

6.3.1. Class Symbol Supplement#

Entities of type a_class_symbol_supplement, pointed to from all class, struct, and union tag symbols, constitute the main repository for front-end-only information about classes.

Among the fields in class symbol supplements are:

  • symbols, the symbol list pointed to from the scope stack entry for the class, retained even after the scope stack entry disappears. The symbols are in declaration order, which is the order in which initialization is to be performed.
  • constructor, a pointer to the sk_member_function or sk_overloaded_function symbol that represents the constructor(s) for the class. As noted above, the constructor symbol cannot be accessed directly by normal name lookup, but only via this pointer.
  • destructor and assignment_operator, pointers to the symbols representing the destructor and operator=() functions for the class. Both could also be accessed by normal name lookup, but having the pointers directly available makes some processing, especially implicit invocations of these functions, more efficient and facilitates deciding whether the compiler needs to generate one.
  • conversion_list, a linked list of entries that provide quick access to the symbols for the user-defined conversion functions for the class (i.e., that define a conversion from the class to some other type). The list includes symbols for inherited conversion functions.
  • routine_fixup_list, a pointer to entities that store cached tokens to allow for delayed scanning of default arguments and the bodies of member functions that are defined inline.

    operator_lookup_namespaces, a pointer to a list of namespaces that are the parent namespaces of the class or one of its base classes. The global scope is explicitly represented on this list by an entry with a NULL namespace pointer. This list is used to specify the namespaces to be searched for operator functions when an overloaded operator function is called with an operand of the class.
  • A number of flags to summarize attributes of the class. Most of these flags are redundant – the attributes could also be computed by examining other data. For instance, any_ref_member is checked when putting out warnings on variables that should be initialized; its value could also be determined by looking at all the nonstatic data members to see if any have reference types.

6.3.2. Class Member Projection Symbols#

Symbols of kind sk_projection provide a mechanism to represent names that are visible in a derived class having been declared in a base class. Each projection symbol has a supplementary entity of type a_projection_descr which identifies the fundamental base class member by pointing to the base class entry and to a symbol from that base class. [5]

Projection symbols are created as needed (in the course of looking up a name) and are added to the scope of the derived class. Once entered in a class scope, they permanently associate the name in the derived class with a declaration in a base class. The next time the name is looked up in the lexical context of the class, the projection symbol will be found directly, without the more expensive search of the base classes.

A projection symbol records a name’s accessibility in the current scope (e.g., it may have been declared public originally but be less accessible in the context of the derived class). It also notes whether a name is ambiguous. These attributes of name inheritance are computed on the initial lookup, when the projection symbol is created, and thus need not be recomputed on subsequent references to the name.

Projection symbols are also used to represent base class member names that are explicitly incorporated into the derived class by means of a using-declaration; for these cases the flag is_using_decl is set to TRUE.

When a using-declaration in a derived class specifies an overloaded member function from a base class, each member of the original overload set will be projected into the base class independently and recorded in an overload set created in the derived class. The latter may contain a mix of sk_member_function symbols from the derived class and sk_projection symbols referring to member functions in one or more base classes. For example:

struct A {
  void f(int);
  void f();
};
struct B {
  void f(double);
};
struct C : public A, public B {
  using A::f;
  using B::f;
  void f(char);
};

Here, the overload set for C::f will include C::f(char) along with projections into C of A::f(int), A::f(), and B::f(double).

On the other hand, when is_using_decl is FALSE in an sk_projection symbol, the projection may refer to an entire overload set and not to any particular instance of the name. In such cases the access value will represent the greatest access available; a particular function may turn out to have less accessibility.

6.3.3. Namespace Projection Symbols#

Symbols of kind sk_namespace_projection are used to represent names made visible by using-declarations that refer to namespace members (using-declarations that refer to class members are described above). As with class member using-declarations, when a using-declaration specifies a overloaded function from another namespace, each member of the overload set is projected into the current scope separately. The resulting overload set may contain a mix of function, function template and namespace projection symbols.

6.4. Template Symbols#

Symbols of kind sk_class_template and sk_function_template have an associated entry of kind a_template_symbol_supplement. Template processing is discussed in detail in another chapter.

6.5. Namespace Symbols#

Symbols of kind sk_namespace have an associated entry of kind a_namespace_symbol_supplement. As described above, the namespace symbol supplement contains the pointers block used by scope stack entries that refer to the namespace. It is possible for more than one scope stack entry for a namespace or namespace extension scope to refer to a single namespace. In such cases the scope stack entries all refer to the one pointers block structure in the namespace symbol supplement. The namespace symbol supplement also contains information used by the name lookup routines when doing lookups in the presence of using-directives.

6.6. Symbol Lookup#

6.6.1. Symbol Locator#

Symbol lookup is the process whereby a name, a qualified name, an operator token, or a conversion result type is used to search for a symbol in the symbol table and, if it exists, return it to the caller. A key intermediary construct in this process is the entity of type a_symbol_locator. It consists of

  • A pointer to a symbol header. (Sometimes error locators have a NULL symbol header pointer; in all other cases a locator will point to a valid symbol header.)
  • A source position, the line and column in the source program at which the reference starts.
  • specific_symbol, a pointer to a symbol. This is usually NULL, but if preliminary processing happens to turn up a symbol (e.g., in validating a qualified name), it is saved here so that the lookup need not be done again.
  • The is_qualified_name, is_global_qualified_name, is_file_scope_qualified_name, and is_super_qualified flags, which specify the kind of qualifier, if any, that preceded the name.
  • is_class_member, a flag that is TRUE for qualified names for which the qualifier names a class. Note that the qualifier may also contain a namespace specifier, but the last qualifier specifies a class name.
  • parent, a union that, for qualified names, contains a pointer to either the class or namespace specified by the qualifier. The is_class_member field is used to specify whether the qualifier is a class or a namespace. This field is also used to specify the class type associated with a vacuous destructor reference.
  • The is_template_id flag is TRUE if the name was followed by a template argument list that has been coalesced. If the name is a class template name, the specific_symbol will point to the template class designated by the template argument list. The do_not_clear_specific_symbol flag is set to TRUE when a class template argument list is coalesced because the template argument list is not retained in the locator, and it would be impossible to recompute the specific_symbol value if it were cleared.
  • template_arg_list is used when an explicit function template argument list is specified, and points to the template argument list that was scanned. The is_template_id flag is also set. An empty template argument list may be specified (e.g., “<>”), so the is_template_id flag should be used to check for the presence of an explicit function template argument list, not the template_arg_list field.
  • The is_operator_name flag and the operator name kind. This flag is TRUE and the associated field is defined only when it is an operator name that is being looked up.
  • The is_conversion_name flag and the conversion result type. This flag is TRUE and the associated field is defined only when it is a user-defined conversion that is being looked up.
  • The is_destructor_name flag is TRUE when the name is a destructor name.
  • The is_vacuous_destructor_reference and is_nonclass_destructor flags, which are used to identify explicit destructor references for classes without destructors or for nonclass types.
  • The is_semivisible_nested_class flag is TRUE when the name is not visible by standard scoping rules but only by a special search that implements the “nested class anachronism” (ARM 18.3.5).
  • The access_control_error_reported flag is TRUE if an accessibility error has already been issued on the symbol.
  • The is_error flag is TRUE when the locator represents an error.
  • The do_not_clear_specific_symbol flag is TRUE if the specific_symbol points to a value that should be retained because it cannot or should not be reproduced by a repeated lookup. This is primarily used for coalesced class template references, as described above, but is also used for specific symbol error locators (error locators that refer to a specially created error symbol).

Symbol locators are necessary because symbol lookup sometimes occurs in several stages at separate points of processing. They provide a way to save and use the intermediate results. For instance, when an identifier is scanned by get_token, the locator is stored in locator_for_curr_id. Often a local copy is made so that the information can be preserved as additional tokens are scanned. Then when it comes time to actually enter the symbol in the symbol table, all the information needed is right there in the locator.

6.6.2. Identifiers and Symbol Headers#

Each name in the source program is looked up as it is scanned. This is necessary in part so that get_token processing can distinguish identifiers from keywords.

find_symbol is called to look up the names in the hash table pointed to by global variable symbol_table. The interface to find_symbol uses a name string with a length rather than a null-terminated string because the string in question is usually part of the source line buffer.

After the bucket number is computed, it searches the list of symbol headers pointed to by the bucket till it finds a matching name. If it finds a match, it moves the symbol header to the front of its bucket list so that frequently-used entries will cluster near the beginning. If no match is found, a new symbol header is created and entered. A pointer to the symbol header (whether newly created or not) is returned in the symbol locator, and a pointer to the list of active symbols associated with the symbol header is also returned.

6.6.3. normal_id_lookup#

find_symbol locates a symbol header but does not actually find the symbol itself. For simple names (e.g., not qualified names) that is done by normal_id_lookup, which searches the symbol header’s active list – and, if appropriate, its inactive list – to find a symbol in the appropriate name space and lexical scope.

normal_id_lookup uses one of two search algorithms. The fast search (which is the only one used in ordinary C mode) simply searches the active list: the first symbol found which meets the search requirements is returned; in C++ mode this search may only be used when there are no scopes on the scope stack that alter the sequence in which scopes are considered for name lookup, and when the lookup is not one of a number of special lookups that disregard certain kinds of symbols.

The second and slower algorithm is required when

  • A class or class reactivation scope is present, in which case one must look for inactive symbols for the class members and for member names that are inherited from base classes.
  • A namespace, namespace extension, or namespace reactivation scope is present, in which case one must look for inactive symbols for the namespace members.
  • A template instantiation scope is present, in which case inactive symbols for the template parameters must be considered, and the scopes must be searched in an irregular order (as described earlier).
  • A using-directive is in effect, in which case one must look for symbols for the namespaces referenced directly or indirectly by the using-directive.
  • A special lookup option has been specified for a given lookup. For example, the lookup may optionally ignore class scopes, etc.

In such cases each scope must be considered individually, since there is no single list that summarizes the visibility of names.

This comes up, for instance, in the non-inline definition of a member function: the class definition has been completed and its names have been removed from the active list, so those names must be searched for on the inactive list. For example,

int i;
struct S { int i; int f(); };
int S::f() { return i; }

Here, when the body of S::f() is being scanned, the symbol for ::i will be on the active list of the symbol header for i, whereas the symbol for S::i will no longer be in scope, having been moved over to the inactive list when the definition of S was completed. However, since the scope for S is reactivated before the member function body is scanned, and since the reactivated scope is examined before the file scope, it is S::i rather than ::i that will be found.

Similarly, when a namespace member is defined outside of its namespace a namespace extension scope is pushed onto the scope stack at the point at which names from the namespace should be visible.

int i;
namespace N { int i; int f(); }
void N::f() { return i; }

Here, as with the class member example shown above, the namespace member N::i is found, hiding ::i.

6.6.3.1. Slow Lookup Processing (scope_stack_lookup)#

scope_stack_lookup and the routines it calls are used to implement the slow lookup. It is passed a range of scope depths to be included in the lookup. Each scope is examined separately. The lookup begins with the scope stack entry indicated by the global variable depth_of_initial_lookup_scope and proceeds by following the previous_scope link in the scope stack entry until a scope entry with no previous scope is found. active_scope_lookup is called for scopes whose symbols are still on the active list and inactive_scope_lookup is called for scopes whose symbols have been transferred to the inactive list. When a class or class reactivation scope is processed, a flag is returned that specifies whether scope_stack_lookup should look for a projection symbol for the class as described below.

When a template instantiation scope has been processed, and a symbol has not yet been found, names from both the context in which the template was defined and from the context in which the template was used must be considered. This is done by calling instantiation_context_lookup, which is described below.

A using-directive can cause names from other namespaces to be visible in a namespace scope (including the global scope) that is currently on the scope stack. The using_directives_apply flag in the scope stack entry is used to indicate that such processing is required. When this flag is set for a namespace scope or the global scope, do_using_directive_lookup is called to search for applicable symbols from namespaces named in using-directives.

The presence of using-directives can result in more than one symbol being found that satisfies the lookup. In such cases do_using_directive_lookup creates a synthesized namespace projection symbol that represents the set of symbols found (See using-Directive Processing and Synthesized Namespace Projection Symbols). The lookup could determine that the name is ambiguous in the current context, in which case a synthesized namespace projection symbol that is flagged as ambiguous is returned.

6.6.3.2. Looking for Class Member Projection Symbols#

When a name is not found in a class, the classes from which it is derived must be searched before the containing scope is examined. This is done by find_projected_symbol, which looks for the name among the base classes (see find_progenitor_symbol and symbol_projected_from_base_class) and if successful creates an sk_projection symbol in the derived class and adds it to either the active or the inactive list of the symbol header for the name. A pointer to the projection symbol, which (as described previously) represents the inheritance, is returned to normal_id_lookup, where it is placed in the locator (field specific_symbol); a pointer to the symbol of which it is the projection is returned from normal_id_lookup as the function return value. Consider this example:

int i;
struct S { int i; };
struct T : public S { int f(); };
int T::f() { return i; }

In this case the scope for T is reactivated when the body of T::f() is scanned. No T::i is found in the scope of T, but before the containing scope (namely the file scope) is searched, the base class must be examined, and so normal_id_lookup calls find_projected_symbol. This search turns up S::i, so again ::i is eclipsed. In addition, a projection symbol for i is created in class T (it points to S::i since it represents the “projection” of i from S into T). Thus, on subsequent searches for i in the scope of T a T::i will be found, namely the projection symbol referring to S::i, and so the base class search will not have to be repeated.

find_projected_symbol and its subroutines also compute the accessibility of the inherited name and determine whether it is ambiguous. No errors are put out at this point, since the name lookup routines are interested in name visibility only; instead, the status is recorded in the projection symbol and diagnostics can be issued later if appropriate.

6.6.3.3. Instantiation Context Lookup#

When a name is looked up as part of a template instantiation, and the name is not found in the local context of the instantiation (the template parameters and the scope(s) associated with the class or function being instantiated) it must be looked up in the instantiation context. The front end implements two different instantiation context lookup algorithms: the one mandated by the standard (referred to as “dependent name lookup”), and the one that existed before dependent name lookup was implemented that more closely reflects existing practice and is more compatible with other compilers.

As described earlier, push_template_instantiation_scope configures the scope stack so that three distinct contexts exist outside of the local context of the instantiation: the template definition scopes, the referencing context scopes, and the scopes common to both.

When doing dependent name lookup, the following steps are used to produce the final result of the lookup:

  • The name is looked up in the context of the template definition and only names visible at the point of the template definition are considered.
  • Only those base classes that do not depend on template parameters are considered.

When not doing dependent name lookup, the following steps are used to produce the final result of the lookup:

  • The name is looked up in both the template definition context and the referencing context by calling scope_stack_lookup with the appropriate starting and ending scope depths. The lookup in the referencing context is suppressed if the lookup in the definition context finds a class member.
  • If none or only one of the lookups produces a symbol, and if the lookup in the definition context did not find a class member, scope_stack_lookup is called again to look in the common scopes.
  • If no symbols were found by the two context lookups, the result of the common scope lookup is used.
  • If only one of the two context lookups produced a symbol, the result of the common scope lookup is used as the value of the context lookup that did not produce a symbol, and the next rule is followed.
  • If the symbol from the referencing context is not a function or function template, it is discarded.
  • If there is still only one symbol from the two context lookups, that symbol is used.
  • If there are two symbols, a synthesized namespace projection symbol is created that represents the result of the lookup.

Note that in both cases, these algorithms describe the result of the “normal” lookup of the name. In addition to the entities made visible by the normal lookup, additional functions and/or function templates may be made visible by argument-dependent lookup.

6.6.3.4. Nested Class Anachronism Lookup#

If the processing described thus far in normal_id_lookup fails to find a qualifying symbol, an additional search implementing the “nested class anachronism” (ARM 18.3.5) is performed by calling find_nested_class_symbol, which finds names of nested classes which, strictly speaking, should not be visible. For example,

struct S {
  struct T {
    int a;
  };
  int b;
};
struct T x;

After the normal search for T is exhausted without turning up a symbol, an additional search is made through the inactive list of the corresponding symbol header for a symbol of a nested class whose innermost non-class containing scope (in this case, the file scope) is still an active scope (in this case, it is). When such a symbol is found, and it is the only symbol that qualifies, it is returned to normal_id_lookup and a warning is issued. [6]

6.6.3.5. Nonreal and Proxy Class Lookup#

If the processing described thus far in normal_id_lookup fails to find a qualifying symbol, and one or more of the classes included in the lookup has a “nonreal” base class, a nonreal class member is created by calling add_member_to_proxy_or_nonreal_class, and the symbol for the newly created class member is returned.

typedef int Y;
template <class T> struct A : T {
  X x;   // assumed to be T::X
  Y y;   // ::Y
};

The nonreal lookup is only done when the normal lookup fails to find a symbol. In the example above, ::Y is found, so no nonreal member is created. The lookup of X fails to find a symbol, so a nonreal member is created.

6.6.3.6. Hide-by-sig lookup#

C++/CLI has a special kind of lookup used for members of managed classes, called “hide-by-sig” lookup. In normal C++ name lookup, a name in a derived class hides all instances of the same name in a base class; this is called “hide-by-name” lookup. In hide-by-sig lookup, on the other hand, a member function in a derived class hides a member in a base class only if their signatures match. A function “voidf(int)” in a derived class, for example, would not hide a function “voidf(char)” in a base class; both would appear in an overload set and might compete with each other. In addition, hide-by-sig lookup considers access – inaccessible members of base classes are not seen by the lookup. And there are special rules regarding interfaces, which are too arcane to describe here. Hide-by-sig lookup only applies to functions.

Hide-by-sig lookup is implemented by building a list of a_hide_by_sig_list_entrys that enumerates all the symbols that co-exist with a given symbol in an overload set. The list is built by use_hide_by_sig_lookup and saved in the symbol’s hide_by_sig_lookup_result field so it only needs to be constructed the first time a hide-by-sig lookup is done on the symbol. Overload resolution traverses that list to build up the set of functions in an overload resolution set. Access is checked while the overload set is built up, and inacessible functions are discarded at that point (different places in the program may have differing access to the base class members; this allows a single general hide-by-sig list to be built and then used with different access contexts).

6.6.3.7. Lookup of Entities Imported from C++/CLI Metadata#

When C++/CLI metadata files are imported by way of a #using directive or implicitly for the entities in mscorlib, only the namespace-scope-level names in the file are added to the symbol table. For a class, for example, the class name is entered as an incomplete type, and its definition and members are not immediately entered. Later, if a lookup is done in the class scope, the members of the class are fetched from metadata and added to the symbol table. This “lazy” lookup makes the use of metadata files much more efficient, since things are not loaded until they are actually needed. This is implemented by piggy-backing on the mechanism that instantiates templates when a member of the template is required: f_instantiate_template_class (often invoked from the macro complete_class_type_is_needed)will call get_definition_of_class to fill in the definition of an entity imported from metadata. This is done automatically whenever a class-scope lookup is done in a class that came from metadata. The definition is loaded, the lookup is then done, and the result is returned to the caller, who has no need to know that the class definition had to be loaded to be able to perform the name lookup.

6.6.3.8. C mode out-of-scope declarations#

When compiling C code, if the processing described thus far in normal_id_lookup fails to find a qualifying symbol, find_out_of_scope_declaration is called to look for an external variable or function that was declared in a scope that is no longer visible. If such a symbol is found, it is reentered in the current scope and returned as the result of the lookup.

void g1(void) {
  extern void f(int);
  f(1);
}
void g2(void) {
  f(2);
}

A symbol for void f(int) is reentered in g2, even though f is no longer visible. This lookup is not performed in strict mode.

6.6.3.9. Lookup Options#

Options flags may be supplied to normal_id_lookup that affect the kinds of symbols that are visible and other aspects of the name lookup process. In certain contexts a name lookup is required to determine whether a given name is a type. In this context a projection symbol should not be created unless the name is a type. The flag IDL_TENTATIVE_TYPE_LOOKUP is used to suppress the creation of a projection symbol in this instance.

A special lookup is also required when scanning a constructor initializer list that may include member names and base class names but may not reference the names of parameters of the constructor. The IDL_SKIP_CURR_FUNCTION_SCOPE flag is used to request this kind of lookup.

IDL_LINKAGE_LOOKUP indicates that the lookup is being used to find a previous declaration of a name and that certain special lookup features should be suppressed and that the lookup should terminate at the nearest enclosing namespace scope. This option suppresses nested class anachronism, nonreal member lookup, and out-of-scope declaration processing.

IDL_SKIP_CLASS_SCOPES causes class and class reactivation scopes to be ignored. This is used when looking up nonmember operator functions.

IDL_INSTANTIATION_CONTEXT is used to indicate that a lookup operation is part of a special instantiation lookup that looks in both the template definition and reference contexts.

6.6.4. Looking Up Names in the Current Scope#

curr_scope_id_lookup looks up a name in the current scope and returns the symbol found or NULL if no symbol is found. The symbol must be a name declared in the scope. using-directives are not considered and projection symbols to base class members are not considered unless the IDL_PROJECTION_SYMBOL_ALLOWED option is specified.

6.6.5. Looking Up Class Members#

class_qualified_id_lookup is used to look up a name in the scope of a class and its base classes. As with normal_id_lookup, the caller provides a symbol locator that contains a pointer to the symbol header for the identifier being looked up. The search begins by looking through the inactive list for a symbol from the scope associated with the class type passed by the caller. If the symbol is not found on the inactive list, the active list is also checked. The active list check is needed when a class-qualified lookup is done while the class is being defined. For example,

struct A {
  struct B {};
  A::B f();  // the A:: is not necessary, but is allowed
};

If the symbol has still not been found, and the lookup is being done in a nonreal class, a symbol for a nonreal class member is created by calling add_member_to_proxy_or_nonreal_class, and the newly created symbol is returned (see Nonreal and Proxy Class Members for more information).

If the symbol has still not been found, the name is checked to see if it is a constructor or destructor for the class. If so, the symbol for the constructor or destructor from the class symbol supplement is returned.

Finally, if the symbol has still not been found, find_projected_symbol is called to search the base classes for a member that may be projected into the class.

6.6.5.1. Nonreal and Proxy Class Members#

A nonreal class is a template class that has one or more template arguments that depend on template parameters. A proxy class is a class that stands in for a template parameter when that template parameter is used in a context that requires a class type. For example:

template <class T> struct A {
  typename T::A      a;
  typename A<T*>::X  b;
  typename T::Z<int> c;
  int n[T::B];
};

In this example a proxy class is created for T so that A and B may be looked up in T. A nonreal class for A<T*> is created, a nonreal template named T::Z is created, and the nonreal instance T::Z<int> is created.

proxy_class_for_template_param creates a proxy class associated with a given template parameter. The proxy class is created the first time that it is needed

create_proxy_or_nonreal_class_member creates a member of a nonreal or proxy class. It is called by add_member_to_proxy_or_nonreal_class which also adds the newly created symbol to the inactive list so that it may be found by subsequent lookups. The kind of entity that is created (type, constant, or template) depends on the kind of lookup that is done, and on whether the “implicit typename” option is enabled. nonreal_member_symbol_kind is used to determine the kind of symbol to be created for a given set of lookup options.

6.6.6. Looking Up Namespace Members#

namespace_qualified_id_lookup is used to look up a name in the scope of a namespace. If the name is not found in the namespace, the name is also sought in any namespaces referenced by using-directives in the specified namespace. lookup_in_namespace is called to look for the name in the specified namespace. If a symbol is found, it is returned; otherwise lookup_in_namespace calls qualified_using_directive_lookup, which loops though the using-directives found in the original namespace and calls lookup_in_namespace recursively to look for the name in those namespaces. qualified_using_directive_lookup always considers all namespaces named in using-directives within a given namespace (i.e., it does not stop searching when a symbol is found in a given namespace). For example, the hierarchy below would be created by the following example:

//   D       E
//    \     /
//     B   C
//      \ /
//       A
namespace E { }
namespace D { }
namespace C { using namespace E; }
namespace B { using namespace D; }
namespace A { using namespace B; using namespace C;}

In such a hierarchy, when looking up a name in namespace A

  • If the name is found in A, the lookup will not look in any other namespace.
  • If the name is not found in A, the lookup will always look in both B and C.
  • If the name is not found in B, the lookup will look in D.
  • If the name is not found in C, the lookup will look in E.

The presence of using-directives can result in more than one symbol being found that satisfies the lookup. In such cases qualified_using_directive_lookup creates a synthesized namespace projection symbol that represents the set of symbols found (see using-Directive Processing and Synthesized Namespace Projection Symbols). The lookup could determine that the name is ambiguous in the current context, in which case a synthesized namespace projection symbol that is flagged as ambiguous is returned.

When a namespace contains an inline namespace (a C++11 feature), or a namespace named in a GNU strong using-diirective, namespace lookup is done slightly differently. Normally, if a name is found in a namespace, other namespaces made visible in that namespace by using-directives are not searched, but a namespace was made visible as an inline namespace (or strong using-directive) is searched even when a symbol was found in the original namespace.

6.6.7. Looking Up Names in the File Scope#

The file scope is also known as the global scope, and, with the addition of namespaces, as the “global namespace”. The same qualified name lookup rules apply to both namespace qualified lookups and file scope qualified lookups. file_scope_id_lookup looks on the active list for a name declared at the file scope. If the name is not found on the active list, the inactive list is seached. If the name is not found in the file scope, qualified_using_directive_lookup is called to look in scopes named by using-directives in the file scope.

In the presence of multiple translation units there can be more than one file scope. The scope pointer for the file scope to be used is passed to file_scope_id_lookup.

6.6.8. Argument-Dependent Lookup#

When an unqualified name is used as the function name in a function call, the name is looked up using both the normal lookup mechanism and also using what is known as “argument-dependent lookup”. Argument-dependent lookup uses the types of the function arguments to produce a set of “associated classes” and “associated namespaces” to be searched for candidate functions.

argument_dependent_lookup is called to perform an argument-dependent lookup. It calls determine_assoc_namespaces_and_classes to determine the associated namespaces and classes to be used for the lookup. Normally, only names from the current translation unit are considered by the argument-dependent lookup process, but when instantiating exported templates (and non-exported templates whose instantiation results from the instantiation of an exported template) argument-dependent lookup is done in each of the translation units on the translation unit stack and the results of the lookups are combined.

Argument-dependent lookup returns a list of symbols representing the candidate functions being found. The symbol list may contain functions, function templates, and overload sets. Argument-dependent lookup is normally done in conjunction with a normal lookup. The symbol representing the normal lookup result is passed to argument_dependent_lookup and is returned as an element of the resulting symbol list. The resulting symbol list may directly or indirectly (as part of an overload set) include multiple instances of a given function or function template. This occurs for several reasons:

  • The symbol list includes the result of both the normal and argument-dependent lookups, which may produce overlapping results.
  • The symbol list includes overload sets which may, as a result of using-directives, include the same functions or function templates.
  • The argument-dependent lookup may be done in multiple translation units when exported templates are used. The same functions or templates may be found in more than one of the translation units.

6.6.9. using-Directive Processing#

using-directives are used to make names from a namespace visible at a certain point in the unqualified lookup process (i.e., the lookup performed by normal_id_lookup). using-directives that appear in namespace scopes can also affect qualified lookups in the namespace in which they appear. Qualified using-directive lookups simply make use of the using-directives list found in the scope entry. Unqualified using-directive lookups, on the other hand, rely on additional data structures that are maintained to optimize the process of finding names made visible by using-directives during the normal_id_lookup process.

A using-directive can appear in a namespace scope (including the files scope), or in a block scope. The scope that contains the using-directive is often not the scope at which the using-directive “applies”, however. A using-directive “applies” at the nearest enclosing scope that is or contains the namespace named in the using-directive, and contains the using-directive. For example,

namespace A {}
namespace B {
  namespace C {
    using namespace A;  // applies to the file scope
  }
  using namespace C;    // applies to B
}
using namespace B::C;   // applies to the file scope

The primary piece of information that is maintained to optimize unqualified using-directive lookups is a field in the namespace symbol supplement that indicates the scope depth of the scope at which a using-directive for that namespace applies, or contains NO_SCOPE_DEPTH if there is no using-directive in effect that makes symbols from the associated namespace visible.

do_using_directive_lookup simply looks through the inactive symbols for a symbol whose parent namespace has a using-directive that applies at the scope depth currently being processed by scope_stack_lookup.

The problem with this kind of state information is that it needs to be reset each time a template instantiation begins and restored each time an instantiation ends. To make the process of clearing and restoring the state information as efficient as possible, the scope stack entry contains a list of “active” using-directives for the scope. An active using-directive is one that appeared in the scope or appeared in one of the namespaces named by a using-directive in the scope.

When a using-directive appears, add_active_using_directive is called to add it to the active using-directive list for the current scope. It then calls add_active_using_directive_to_scope, which actually adds it to the list.

using-directives are transitive, meaning that any using-directives that appear in the namespace nominated by a using-directive are treated as if they also appeared in using-directives in the scope containing the original using-directive. add_active_using_directives_for_namespace is called to add the namespaces that appear in using-directives in the nominated namespace to the active using-directive list, which it does by recursively calling add_active_using_directive_to_scope. When add_active_using_directive adds a using-directive to a namespace scope it searches through the scope stack to find active using-directive lists that refer to the namespace containing the new using-directive. If such scopes are found, add_active_using_directive_to_scope is called to add the newly nominated namespace to those scopes.

add_active_using_directive_to_scope determines the scope depth at which a using-directive applies. This scope depth is recorded in the namespace symbol supplement of the namespace being added. The scope stack entry associated with that scope depth is updated to indicate that there are using-directives that apply when that scope is reached in an unqualified lookup.

set_active_using_list_scope_depths is called to reset and restore the state of the scope depth information in the namespace symbol supplement and the using_directives_apply field in the scope stack entry. This is done at the beginning and end of a template instantiation scope, and is also done whenever a scope containing active using-directives is popped off of the stack.

6.6.10. Synthesized Namespace Projection Symbols#

The addition of namespaces creates situations in which the result of a lookup can be a set of symbols from different scopes. A synthesized namespace projection symbol is created to represent the result of such a lookup. Synthesized namespace projection symbols are created

  • by an unqualified lookup in which the symbols found were made visible by a using-directive,
  • by an qualified lookups in which the symbols found were made visible by a using-directive, and
  • by an instantiation context lookup in which symbols were found in both the definition and referencing contexts.

The synthesized_namespace_projection flag in the symbol entry identifies a symbol as a synthesized namespace projection. Such symbols can be of kind sk_namespace_projection or sk_overloaded_function. In the latter case, each of the entries on the list of overloaded functions points to a sk_namespace_projection symbol.

Where synthesized namespace projection symbols are used, the lookup must result in at most one symbol being found, or, all of the symbols found must be functions that can be combined into an overload set. If more than one symbol is found and they are not all functions, the lookup is ambiguous. The ambiguous flag in the symbol entry is set in that case.

The following steps are used to construct a synthesized namespace projection symbol for a given lookup:

  1. find_synthesized_projection_symbol is called to find a previously created namespace projection symbol. This is done so that the space used by a previously created synthesized projection symbol can be reused. The actual information pointed to by the symbol is not reused because it may have changed since the last lookup was done. More specifically, a set of overloaded functions is assumed always to contain at least the functions that were previously in the set. When looking for a previously created symbol, the lookup options used for the new lookup must match the options used when the symbol was created. A number of symbols with the same name but with different lookup options may be created in a given scope. Some combinations of lookup options are considered to be sufficiently unusual that they are not entered in the symbol header for reuse. Such symbols are identified by setting the do_not_reuse flag in the symbol entry. find_synthesized_projection_symbol suppresses the search and returns NULL when asked to find a symbol with a set of lookup options for which symbols cannot be reused. The lookup options used when a synthesized namespace projection symbol is created are stored in a set of flags in the symbol entry. The flags used for this purpose are

    • qualified_lookup
    • must_be_class_or_namespace_lookup
    • must_be_tag_lookup
    • tentative_type_lookup
    • do_not_reuse
    • instantiation_context_lookup
  2. As symbols are found that match the lookup criteria, add_symbol_to_lookup_set is called with a pointer to the previously created synthesized projection symbol (if any) and the symbol to be added. The symbol returned is used as the new synthesized projection symbol.
  3. add_symbol_to_lookup_set calls merge_function_into_lookup_set when the symbol being added is a function or function template. Before a function is actually added, the current lookup set is examined, by calling already_in_lookup_set, to make sure it is not already present.
  4. add_symbol_to_lookup_set determines whether the symbol being added can coexist with any symbols already in the lookup set, whether the new symbol should hide a previous symbol, or whether the lookup is ambiguous.

Reusable synthesized projection symbols are entered on the “other” symbols list in the symbol header and on the synth_namespace_projection_symbols list in the pointers block for the scope with which they are associated. Symbols that are not reusable are only entered on the list in the pointers block. Synthesized projection symbols that are created by a qualified lookup are entered in the scope of the namespace of the qualifier. All other synthesized projection symbols are entered into the current scope.

6.6.11. Overloaded Operators#

Overloaded operators are not looked up by name, so find_symbol cannot be used to locate the symbol header. Instead, the header is found by indexing into the array of symbol header pointers opname_symbol_table. The index value will be a member of the enumeration an_opname_kind, which maps to the subset of the enumeration a_token_kind that comprises those operators subject to overloading. [7]

When an overloaded operator is explicitly named (for instance, operator+ or A::operator new), it is represented by a tok_identifier “pseudo-token” somewhat like qualified names. The symbol locator’s is_operator_name flag will be set and the field opname will identify the operator being overloaded. make_opname_locator is called to set these fields and as well to create and enter the symbol header if one does not yet exist. Scanning the token sequence representing an overloaded operator is handled by f_get_opname , [8] which does error checking and also deals with conversion operators (discussed below).

nonmember_operator_function_lookup is used to find the set of nonmember overloaded operator functions for a given operator kind and a pair of operand types. When looking for a nonmember operator function, one must do a normal (i.e, unqualified) lookup of the operator name, and one must also look up the operator name in the namespaces of the operand types and the base classes of the operand types. For each class type, there is a list of namespaces in which to look for nonmember operator functions. The list is pointed to by the operator_lookup_namespaces field of the class symbol supplement. The search is done by looking at both the active and inactive lists from the appropriate symbol header and checking whether the parent namespace of the symbol is on the operator lookup namespaces list of one of the operand class types. Note that only immediate members of the namespaces on the operator lookup namespaces list are considered. using-directives in those namespaces have no effect on the lookup. The result of the lookup is a pointer to a list of entry of kind a_symbol_list_entry or NULL if no such operator functions have been declared. Each entry points to a symbol that meets the lookup criteria. The caller is responsible for freeing the symbol list when it is no longer needed by calling free_list_of_symbol_list_entries. opname_function_symbol is a simplified version of nonmember_operator_function_lookup that is used for looking up new and delete operators, which cannot be declared in namespaces.

global_operator_new_or_delete_symbol looks up symbols for the default versions of operator new and delete at file scope (i.e., ::operator new(size_t) and ::operator delete(void*)), but unlike opname_function_symbol it will, if no such symbol exists yet, implicitly declare the function, creating and entering its symbol and adding the routine entry to the IL.

6.6.11.1. Template Conversion Operators#

When a class contains a template conversion operator, the lookup of a particular conversion operator may require the partial instantiation of the conversion operator template. For example:

struct A {
  template <class T> operator T();
};
int main(){
  A a;
  a.operator int();
}

Class A does not contain an operator int conversion function that can be found by the normal lookup mechanism. Instead, the lookup operation must match the required type (int) with the available conversion templates to see if a match can be found. look_up_conversion_template_instance is called to look for the matching conversion template. If more than one matching template is found a symbol that is marked as ambiguous is returned.

If the result type of the conversion depends on a template parameter, or if the class type on which the conversion operator is being called is a nonreal class, it is not possible to find a matching conversion operator. In such cases an “unknown conversion function” entry is created. This is represented by a special kind of template parameter constant that contains a pointer to the conversion result type.

6.6.12. Constructors and Destructors#

Several routines search for a constructor or destructor of a particular kind in a given class.

select_default_constructor finds the default constructor for a given class. A default constructor is one that can be called with no arguments. If the class has no default constructor, an error is issued.

select_destructor finds the destructor for a given class. If the class has no destructor, it returns NULL.

select_copy_constructor finds a copy constructor for a given class. It can be told to look for a copy constructor that will accept const- or volatile-qualified objects. If no appropriate copy constructor exists, an error is issued.

6.6.13. User-Defined Conversions#

Like overloaded operators user-defined conversion functions (or “conversion operators”) cannot be looked up by name. Rather, the symbol header is looked up by result type, using the linked list of entries of type a_conversion_header and pointed to by conversion_header_list, as described earlier. Since all user-defined conversion functions are member functions, once the symbol header is found, searching for a symbol simply involves discriminating on the basis of the parent.class_type field.

An explicitly named conversion function (e.g., operator int()), is represented (like overloaded operators) by an identifier token; in its associated symbol locator is_conversion_name is TRUE and conversion_result_type points to a type entry for the result of the conversion. The routine called to find the conversion header (and thereby the symbol header) and to set these fields is make_type_conversion_locator. It also creates the conversion header (as well as the symbol header) if none exists yet. The token sequence representing a conversion function is scanned by scan_conversion_operator, which is called from f_get_opname.

Sometimes the question to be answered in processing an expression is, “What user-defined conversion function exists to convert an object of class A to type T?” In such cases it would be straightforward to find the appropriate symbol by using the result type to identify the conversion header and then the class type to choose among the associated conversion function symbols. But at other times – for instance, when looking for the best implicit conversion or when there are inherited conversion functions – the question is subtly different: “Which, if any, of the user-defined conversion functions that exist for class A is most suitable for converting to type T?” To optimize answering this question, another construct appears in the symbol table, a_conversion_list_entry, each instance of which points to a symbol entry representing a conversion function. A class’s symbol supplement contains a pointer to such a list, which is updated in decl_member_function each time a conversion function is declared; the list also contains projection symbols to represent inherited conversion functions (see project_base_class_conversion_functions).

6.7. “Extern” Symbols#

When it is required that distinct declarations in different name scopes refer to the same entity, an extra symbol is created to represent the shared entity. For instance,

void f() {
  extern int i;
}
void g() {
  extern int i;
}
extern int i;

In this example, although all three declarations of i refer to the same variable, three independent sk_variable symbols are created to represent it, one in the scope of function f, one in the scope of function g, and one in the file scope. So a fourth symbol is created at file scope to represent i – its kind is sk_extern_variable, and it is created at the first declaration of i. Then, even after the sk_variable symbol for the declaration of i in function f goes out of scope, the sk_extern_variable symbol is still around to be found again on the subsequent declarations of i; this assures that all three i symbols point to the same IL entity. A similar technique, using sk_extern_function symbols, is employed for functions. These “extern” symbols are put on the “other” symbols list instead of being put on either the active or inactive list and so are not found during normal name lookup.

The routine that manages all this is find_external_symbol. It also has special logic, subject to configurable control, to deal with instances in which different source names resolve to the same “external” name because of object language constraints – e.g., where linkers can handle only so many characters or are not case sensitive.

6.8. Symbol Management#

It remains to describe briefly some of the other low-level routines and macros used to manage the symbol table and return information about symbols.

  • alloc_symbol is called to allocate a symbol entry of a specified kind. When a NULL symbol header is passed, it allocates an error symbol, one whose header is error_symbol_header. It calls clear_symbol and set_symbol_kind to initialize the new entities; the latter also allocates the supplementary constructs for class and projection symbols. Other routines called to allocate and initialize the entities from which the symbol table is constructed are
    • alloc_symbol_header,
    • alloc_conversion_header, and
    • alloc_conversion_list_entry
  • enter_namespace_projection_symbol is called to create new namespace projection symbols and enter them in the symbol table. Passed a pointer to the fundamental symbol, it allocates and initializes the symbol and links it into the symbol table and scope list. This routine is needed because a namespace projection symbol cannot be entered into the symbol table via enter_symbol because that doesn’t allow specification/setting of the fundamental symbol and the fundamental symbol is used by enter_symbol for checking compatibility with previous declarations of the name in the current scope.
  • enter_synthesized_namespace_projection_symbol creates new synthesized namespace projection symbols and enters them on the “other” symbols list if they can be reused. This routine also determines the scope list on which the symbol should be placed based on the lookup options and the qualifier namespace, and sets the lookup option fields in the symbol entry.
  • enter_symbol is called to create new symbols and enter them in the symbol table. Passed a symbol kind, it allocates and initializes the symbol, marks it as having been referenced, and links it onto the symbol header’s active list and onto the current scope’s symbol list.
  • reenter_symbol takes an existing symbol (one that may have been entered already and then removed) and links it into the symbol table.
  • full_enter_symbol calls find_symbol to locate the symbol header for a name and then calls enter_symbol. This is convenient when creating symbols for the reserved names of keywords and predefined macros (see enter_keyword and enter_predef_macro in fe_init.c).
  • link_symbol_into_symbol_table, called from enter_symbol and elsewhere, links a symbol onto the symbol header’s active list. Though it normally ends up adding the symbol to the front of the list, it is careful to account for special cases by determining the precise location based on the scope number and name space overloading. It will issue a redeclaration error when it detects that the name has already been declared in the current scope.
  • add_symbol_to_scope_list is also called from enter_symbol; it adds a symbol to the end of the symbol list [9] of the scope stack entry at the specified scope depth.
  • enter_overloaded_symbol is called to enter a function symbol in the symbol table when its name is overloaded. First, it allocates and initializes the sk_routine or sk_member_function symbol. Then, if an sk_overloaded_function symbol already exists for the name, the new symbol is added to its list. If not, a new sk_overloaded_function symbol must also be created; it will replace the existing function symbol in the symbol table.
  • make_namespace_projection_symbol allocates and initializes namespace projection symbols. A pointer to the fundamental symbol is provided.
  • make_projection_symbol allocates and initializes sk_projection symbols, leaving it to the caller, find_projected_symbol, to add the new symbol to the symbol table, since it may be inserted in either the active list or the inactive list of the symbol header.
  • make_template_class_symbol creates and initializes a symbol for an instance of a class template. It is linked in the template’s instantiation list rather than added directly to the symbol table.
  • make_unnamed_tag_symbol is invoked to allocate and initialize a tagless class, struct, and union symbol. It is not actually entered into the symbol table since, like error symbols, it uses a special symbol header (see static variable unnamed_tag_symbol_header). If the class subsequently acquires a name, relink_unnamed_tag_symbol is called to modify the symbol entry and add it to the symbol table.
  • unlink_symbol_from_symbol_table is called from pop_scope for each name that passes out of scope: it removes the symbol from the symbol header’s active symbol list, but the symbol remains on the scope list.
  • remove_symbol removes a symbol from the symbol table, e.g., for a macro definition when a #undef is encountered. unlink_symbol_from_symbol_table is called to remove a symbol from the active list and remove_symbol_from_scope_list is called to remove it from the symbol list of the scope stack entry.
  • When cross-reference information is being generated, write_xref_entry is called to create a record describing a reference and to write it to the file indicated by f_xref_info.
  • record_symbol_declaration is called for every declaration and definition of a name. It sets the source sequence number if this is the initial declaration, may set the source position (for initial declarations and for definitions), optionally calls write_xref_entry, and (when GENERATE_SOURCE_SEQUENCE_LISTS is TRUE) calls sym_update_source_sequence_list.
  • mark_defined and mark_declared are macro interfaces to record_symbol_declaration, for definitions and non-defining declarations, respectively.
  • record_symbol_reference is called for nondeclarative references to symbols. It optionally calls write_xref_entry, specifying the kind of reference (see a_symbol_reference_kind), sets the symbol’s referenced flag, and may set the referenced flag in the IL entry. For variables, it checks for the need to issue a “used-before-set” warning and calls mark_variable_value_set if appropriate.
  • mark_variable_value_set is called to set the value_has_been_set flag in the symbol of a variable or static data member. For a variable entry that represents function parameter or a handler parameter it also sets param_value_has_been_changed.
  • sym_update_source_sequence_list is called (when GENERATE_SOURCE_SEQUENCE_LISTS is TRUE) to generate a source sequence entry for a declaration. (Source sequence lists are discussed in greater detail in the context of the intermediate language.)
  • clear_locator is a macro used to clear the fields of a locator and initialize source_position to a specified value; set_to_error_locator is similar but is used when errors are encountered.
  • set_to_named_error_locator turns a locator into an error locator – but one that preserves information about the identifier with which the locator is associated.
  • tildize_locator takes the locator for an identifier token, turns the name into a destructor name by prepending a ~, calls find_symbol to find (or create) the associated symbol header, and returns a locator for new name.
  • change_class_locator_into_constructor_locator takes a locator for a class name (its specific_symbol field will point to the class, struct, or union symbol) and modifies it to point to a separate header of the same name that serves as the header for a constructor. [10]
  • make_opname_locator creates a locator to represent a token sequence consisting of operator and the token (or token pair) for an overloadable operator. It looks for the corresponding header in the opname_symbol_header array; if none is found, it creates one and adds it to the array.
  • make_type_conversion_locator creates a locator to represent a token sequence consisting of operator and tokens making up a type name. It looks for the corresponding header by scanning conversion_header_list; if none is found, it creates one, along with a new conversion header entry, and adds the latter to the head of the linked list.
  • make_locator_for_symbol is called to initialize a locator based on an already existing symbol.
  • set_source_corresp is called to establish a correspondence between an IL entry and a symbol. It links the IL entry back to the symbol.
  • set_decl_sequence_number is called to set a declaration sequence number for a symbol. Such numbers identify the sequential position of one declaration relative to others in the same translation unit. When a name is multiply declared, the position of its definition is recorded in the declaration sequence number.
  • A number of macros in symbol_tbl.h provide the preferred method of identifying a given symbol’s special characteristics:
    • is_class_symbol
    • is_tag_symbol
    • is_type_symbol
    • is_function_symbol
    • is_member_function_symbol
    • is_constructor_symbol
    • is_destructor_symbol
    • is_copy_constructor_symbol
  • type_symbol_type is a macro that may be invoked for symbols for which is_type_symbol is TRUE; it returns a pointer to the associated type entry.
  • symbol_supplement_for_class takes a type entry of kind tk_class, tk_struct, or tk_union and returns a pointer to the associated symbol’s class symbol supplement entry.

6.9. Access Control#

The function of access control is to determine whether or not a given name can be accessed from the current location in the source program. Every reference to a name, whether qualified or not, whether part of an expression or a type, is checked to make sure the identifier is unambiguous and accessible. There are several things to remember about access control:

  • Only class members are subject to access control. Non-class members are always accessible.
  • Access control is site-specific. A given member can be accessible from one place in a program and not accessible from somewhere else. In particular, members and friends of a class have special access to the members of the class and its base classes.
  • Access control limits access, not visibility. Members that are not accessible are still visible. They are found, the access error is issued, and then (the way the front end does it) they are used anyway.
  • When members are inherited into derived classes, the access to the members in the derived class is affected by the type of derivation of the derived class (see ARM 11.2).
  • The same member may have different access in different classes, and may be accessible at a given point in the program through one qualified name and not through another. In other words, access control applies to names rather than objects.
  • A name is ambiguous if the name refers to two or more entities, with no reason to prefer one over the others. Only inherited class members can be ambiguous. Ambiguity checking is done before access control checking.

The problem of access control, then, reduces to “given a symbol for a class member, which might be a projection symbol, is the member unambiguous and do we have access to it from the current location in the program?”

The ambiguity aspect is easily dealt with: the ambiguous flag in the symbol indicates whether or not it is ambiguous. The flag can be set only in projection symbols and in symbols for sets of overloaded functions synthesized for lookups done in the presence of using-directives.

Access checking is harder.

The access to a class member, in its simplest form, is given by the source_corresp.access field in the member’s IL entry.

When the member is a direct member of a class, there will be a symbol entry pointing directly to the IL entry, and the access to that symbol is simply the access indicated in the IL entry. When the member is inherited, however, each derived class will have a projection symbol that points to the fundamental symbol. The access to the member through a projection symbol is filtered through the class derivations that lead to the derived class containing the projection symbol; that access is indicated in the projection symbol.

If the simple access in the symbol or in the projection symbol is as_public, then the member is accessible. Otherwise, some kind of member access privilege is required to access it.

Member access privileges are available when the current position is inside a class or a function. Being inside a class confers special access to members of the class and to classes that have befriended the class. Being inside a function confers special access to classes that have befriended the function, and, if the function is a member function, to the class of which it is a member (and classes that have befriended it).

The scope_stack provides the necessary information about the functions and/or classes that the current position is inside of.

For there to be special access to a member x across the derivation represented by a projection symbol, there must exist a class A somewhere on the derivation (including either of the endpoints) such that:

  1. A is accessible (meaning that its public members are accessible) from the current position in the program, starting from the most-derived class on the derivation. In other words, we can implicitly cast a pointer to the derived class we start with to a pointer to A.

  2. We have member access to A.

  3. x is accessible in A, meaning that if one starts with the access to x in the fundamental class and works outward through the derivation steps until one gets to A, there is some remaining access at that point.

The implementation of the above is broken up as follows:

  • access_for_symbol extracts the simple access for a symbol.
  • access_to_end_of_path determines the effect of a set of derivation steps on such an access.
  • have_access_to_symbol determines whether or not there is access to a symbol. Most of the work is done in have_access_across_path.
  • The macro is_accessible_imm_base_class determines whether or not a base class is accessible (ARM 4.6). The macro is also used directly in determining whether or not base class casts and the like are valid. The function is_accessible_base_class is similar, but is not limited to immediate base classes.
  • have_member_access_privilege determines whether or not there is member access privilege to a given class.

    When there are protected members or protected derivation steps involved, derived classes can have a limited form of member access (see have_protected_member_access_privilege).

    The additional restrictions on access to protected members of ARM 11.5 require special processing in the expression routines over and above the access checking done here. The checking there ends up calling the macro check_protected_member_access, which in turn calls the function f_check_protected_member_access.
  • check_ambiguity_and_verify_access (a macro) is the main interface to access checking. It in turn calls member_check_ambiguity_and_verify_access for class members.

Overloaded functions pose some special problems, because they are represented as several symbols (which may have distinct accesses) clustered under one overloaded function symbol. When the cluster is inherited into a derived class, a single projection symbol points to the overloaded function symbol. This means that:

  • One cannot determine the access to a name in a set of overloaded functions until one has determined which specific function is being referenced. For that reason, when an overloaded function symbol is passed through the normal access checking routines, it is always judged to be accessible. Note that ambiguity checking can still be done in the usual way.
  • Once one has determined which function is being referenced, one cannot use the normal interface to the access checking routines, because the projection symbol does not point to the specific function.

    overload_check_ambiguity_and_verify_access does the access check in this case on the basis of a projection symbol and a (not directly attached) specific function symbol. In the presence of member using-declarations, the specific function symbol can be a projection symbol.

It is not possible to determine the access to certain names used in namespace scope declarations at the point at which the name is seen. In the following example, the accessibility of A::B in the declaration of A::f cannot be determined until we know we are scanning a declaration of a member of class A. In the friend case, we cannot determine the access to A::B until we know the signature of the function being declared, because we have to determine that it has been befriended by class A.

class A {
  class B {};
  B f();
  friend B g(B);
};
A::B A::f(){ B b; return b; }
A::B g(A::B b){ return b; }

When a namespace scope declaration is processed, a flag is set in the scope stack indicating that any access errors that occur should, instead of being issued, be recorded in a list pointed to by the scope stack entry. Later, when more information about the declaration is available, the accessibility is rechecked. If the name is now accessible, the entry is removed from the list of deferred access checks. If it is still not accessible, it will either remain on the list if the defer access checks flag is still set, otherwise the error will be issued.

begin_deferral_of_access_checks and end_deferral_of_access_checks are macros used to bracket the processing of a declaration that requires deferred access checking.

perform_deferred_access_checks is used to retry the access checks that had failed earlier, and is called during declarator processing to check the accessibility of member declarations. This is done after the class reactivation scope for the member declaration has been pushed.

For the accessibility of friend declarations to be checked, the scope stack must contain information about the function that is being declared. The function access scope fulfills this need. After a complete function declaration has been scanned, perform_deferred_access_checks_for_function is called with a pointer to the routine that has been declared. It pushes a function access scope and then calls perform_deferred_access_checks to determine whether the names are now accessible.