0% found this document useful (0 votes)

29 views51 pages

CD Unit 3 - Merged

Uploaded by

tr.anuvarshini

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

29 views51 pages

CD Unit 3 - Merged

Uploaded by

tr.anuvarshini

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

You are on page 1/ 51

Unit - 3

SYNTAX DIRECTED TRANSLATION & INTERMEDIATE CODE GENERATION

Syntax directed translation:

Syntax directed definitions, Bottom- up evaluation of S-
attributed definitions, L- attributed definitions, Top-
down translation, Bottom-up evaluation of inherited
attributes.
Type Checking:
Type systems, Specification of a simple type checker.

1
NEED FOR SEMANTIC ANALYSIS
 Semantic analysis is a phase by a compiler that adds semantic information to the parse tree
and performs certain checks based on this information.
 It logically follows the parsing phase, in which the parse tree is generated, and logically
precedes the code generation phase, in which (intermediate/target) code is generated. (In a
compiler implementation, it may be possible to fold different phases into one pass.)
 Typical examples of semantic information that is added and checked is typing information
( type checking ) and the binding of variables and function names to their definitions
( object binding ).
 Sometimes also some early code optimization is done in this phase. For this phase the
compiler usually maintains symbol tables in which it stores what each symbol (variable
names, function names, etc.) refers to.

Following things are done in Semantic Analysis:

1. Disambiguate Overloaded operators : If an operator is overloaded, one would like to specify

the meaning of that particular operator because from one will go into code generation phase next.

2. Type checking : The process of verifying and enforcing the constraints of types is called type checking.

 This may occur either at compile-time (a static check) or run-time (` dynamic check).
 Static type checking is a primary task of the semantic analysis carried out by a compiler.
 If type rules are enforced strongly (that is, generally allowing only those automatic type conversions
which do not lose information), the process is called strongly typed, if not, weakly typed.

3. Uniqueness checking : Whether a variable name is unique or not, in the its scope.

4. Type coercion : If some kind of mixing of types is allowed. Done in languages which are not
strongly typed. This can be done dynamically as well as statically.

5. Name Checks : Check whether any variable has a name which is not allowed. Ex. Name is same
as an identifier( Ex. int in java).

2
 A parser has its own limitations in catching program errors related to semantics, something
that is deeper than syntax analysis.
 Typical features of semantic analysis cannot be modeled using context free grammar
formalism.
 If one tries to incorporate those features in the definition of a language then that language
doesn't remain context free anymore.
 These are a couple of examples which tell us that typically what a compiler has to do
beyond syntax analysis.
 An identifier x can be declared in two separate functions in the program, once of the type
int and then of the type char. Hence the same identifier will have to be bound to these two
different properties in the two different contexts.

Semantic Errors
We have mentioned some of the semantics errors that the semantic analyzer is expected to
recognize:

 Type mismatch
 Undeclared variable
 Reserved identifier misuse.
 Multiple declaration of variable in a scope.
 Accessing an out of scope variable.
 Actual and formal parameter mismatch.
Syntax Directed Translation
The Principle of Syntax Directed Translation states that the meaning of an input sentence is related
to its syntactic structure, i.e., to its Parse-Tree. By Syntax Directed Translations we indicate those
formalisms for specifying translations for programming language constructs guided by context-
free grammars.
– We associate Attributes to the grammar symbols representing the language constructs.
– Values for attributes are computed by Semantic Rules associated with grammar
productions.
Evaluation of Semantic Rules may:
– Generate Code;
– Insert information into the Symbol Table;

3
– Perform Semantic Check;
– Issue error messages;
– etc.
 There are two ways to represent the semantic rules associated with grammar symbols.
1. Syntax-Directed Definitions (SDD)
2. Syntax-Directed Translation Schemes (SDT)

Syntax Directed Deﬁnitions

Syntax Directed Definitions are a generalization of context-free grammars in which:
1. Grammar symbols have an associated set of Attributes;
2. Productions are associated with Semantic Rules for computing the values of attributes.
• Such formalism generates Annotated Parse-Trees where each node of the tree is a record with a
field for each attribute (e.g., X.a indicates the attribute a of the grammar symbol X).
 The value of an attribute of a grammar symbol at a given parse-tree node is defined by a
semantic rule associated with the production used at that node.

ATTRIBUTE GRAMMAR
 Attributes are properties associated with grammar symbols. Attributes can be numbers,
strings, memory locations, data types, etc.
 Attribute grammar is a special form of context-free grammar where some additional
information (attributes) are appended to one or more of its non-terminals in order to
provide context-sensitive information.

 Attribute grammar is a medium to provide semantics to the context-free grammar and it

can help specify the syntax and semantics of a programming language. Attribute grammar
(when viewed as a parse-tree) can pass values or information among the nodes of a tree.

Ex: E → E + T { E.value = E.value + T.value }

 The right part of the CFG contains the semantic rules that specify how the grammar should
be interpreted. Here, the values of non-terminals E and T are added together and the result
is copied to the non-terminal E.

 Semantic attributes may be assigned to their values from their domain at the time of
parsing and evaluated at the time of assignment or conditions.

4
 Based on the way the attributes get their values, they can be broadly divided into two
categories : synthesized attributes and inherited attributes
ATTRIBUTES

Synthesized Inherited
1. Synthesized Attributes: These are those attributes which get their values from their
children nodes i.e. value of synthesized attribute at node is computed from the values of
attributes at children nodes in parse tree.
 To illustrate, assume the following production:
EXAMPLE: S -> ABC
S.a= A.a,B.a,C.a
If S is taking values from its child nodes (A, B, C), then it is said to be a synthesized attribute, as
the values of ABC are synthesized to S.
Computation of Synthesized Attributes

 Write the SDD using appropriate semantic rules for each production in given grammar.
 The annotated parse tree is generated and attribute values are computed in bottom up
manner.
 The value obtained at root node is the final output.
Consider the following grammar:
S --> E
E --> E1 +
T E --> T
T --> T1 *
F T --> F
F --> digit

The SDD for the above grammar can be written as follow

PRODUCTIONS SEMANTIC RULES

5
Let us assume an input string 4 * 5 + 6 for computing synthesized attributes. The annotated parse
tree for the input string is
S

 For computation of attributes we start from leftmost bottom node. The rule F –> digit is
used to reduce digit to F and the value of digit is obtained from lexical analyzer which
becomes value of F i.e. from semantic action F.val = digit.lexval.
 Hence, F.val = 4 and since T is parent node of F so, we get T.val = 4 from semantic action
T.val = F.val.
 Then, for T –> T1 * F production, the corresponding semantic action is T.val = T1.val *
F.val . Hence, T.val = 4 * 5 = 20
 Similarly, combination of E1.val + T.val becomes E.val i.e. E.val = E1.val + T.val = 26.
Then, the production S –> E is applied to reduce E.val = 26 and semantic action associated
with it prints the result E.val . Hence, the output will be 26.

ANNOTATED PARSE TREE

 The parse tree containing the values of attributes at each node for given input string is
called annotated or decorated parse tree.
2. Inherited Attributes: These are the attributes which inherit their values from their parent or
sibling nodes. i.e. value of inherited attributes are computed by value of parent or sibling nodes.
6
EXAMPLE:

A --> BCD { C.in = A.in, C.type = B.type }

B can get values from A, C and D. C can take values from A, B, and D. Likewise, D can
take values from A, B, and C.

Computation of Inherited Attributes

 Construct the SDD using semantic actions.
 The annotated parse tree is generated and attribute values are computed in top down
manner.
Consider the following grammar:

D --> T L
T --> int
T --> float
T --> double
L --> L1, id
L --> id

The SDD for the above grammar can be written as follow

PRODUCTIONS SEMANTIC RULES

DTL L.in = =T.type

7
 Let us assume an input string int a, c for computing inherited attributes. The annotated parse
tree for the input string is

 The value of L nodes is obtained from T.type (sibling) which is basically lexical value
obtained as int, float or double.
 Then L node gives type of identifiers a and c. The computation of type is done in top
down manner or preorder traversal.
 Using function Enter_type the type of identifiers a and c is inserted in symbol table at
corresponding id.entry.

Implementing Syntax Directed Deﬁnitions

1. Dependency Graphs
 Implementing a Syntax Directed Definition consists primarily in finding an order for the
evaluation of attributes
 Each attribute value must be available when a computation is performed.
 Dependency Graphs are the most general technique used to evaluate syntax directed
definitions with both synthesized and inherited attributes.
 Annotated parse tree shows the values of attributes, dependency graph helps to determine
how those values are computed

8
 The interdependencies among the attributes of the various nodes of a parse-tree can be
depicted by a directed graph called a dependency graph.
 There is a node for each attribute;
 If attribute b depends on an attribute c there is a link from the node for c to
the node for b (b ← c).
 Dependency Rule: If an attribute b depends from an attribute c, then we need to ﬁnd the
semantic rule for c ﬁrst and then the semantic rule for b.

Dependency graph for declaration statements.

9
2. Evaluation order
 A dependency graph characterizes the possible order in which we can calculate the
attributes at various nodes of the parse tree.
 If there is an edge from node M to N, then the attribute corresponding to M first be
evaluated before evaluating N.
 Thus the allowable orders of evaluation are N1,N2…..,Nk such that if there is an edge
from Ni toNj then i<j
 Such an ordering embeds a directed graph into a linear order, and is called a topological
sort of the graph.
 If there is any cycle in the graph ,then there is no topologicalsort.ie, there is no way to
evaluate SDD on this parse tree.

TYPES OF SDT’S

1. S –attributed definition
2. L –attributed definition

S-attributed definition

 S stands for synthesized

 If an SDT uses only synthesized attributes, it is called as S-attributed SDT.
EXAMPLE:
ABC { A.a=B.a,C.a}
 S-attributed SDTs are evaluated in bottom-up parsing, as the values of the parent nodes
depend upon the values of the child nodes.
 Semantic actions are placed in rightmost place of RHS.
ABCD{ }.
 Note: (Also write SDD for desk calculator as example).

10
L –attributed definition
 L stands for one parse from left to right.

 Ie, If an SDT uses both synthesized attributes and inherited attributes with a restriction
that inherited attribute can inherit values from parent and left siblings only, it is called as
L-attributed SDT.
EXAMPLE:
A BCD {B.a=A.a, C.a=B.a}
C.a=D.a  This is not possible
 Attributes in L-attributed SDTs are evaluated by depth-first and left-to-right parsing
manner.
 Semantic actions are placed anywhere in RHS.
A{ }BC
B{ }C
BC{ }
 Note: (Also write SDD for declaration statement as example)
If an attribute is S attributed, it is also L attributed.
Evaluation of L-attributed SDD

11
SDD for desk calculator/SDD for evaluation of expressions

 Evaluate the expression 3*5+4n using the above SDD both in bottom up and top
down approach
Solution: Bottom up evaluation for this expression is shown below
 In both case first we need to draw the parse tree.

12
 Then traverse from top to down left to right
 In bottom up approach, whenever there is a reduction ,go to the production and carry out
the action

Solution: Top down approach:

 First we need to draw the parse tree.

 Then traverse from top to down left to right
 In Top down up approach, whenever we come across a semantic action, it is performed.

13
Workout problems

1. Evaluate the expression 4-6/2

2. Evaluate the expression 2+ (3*4)

Traversal

Procedure visit (n:node)

begin

for each child m of n, from left to right do

visit(m)

14
evaluate semantic rule at node n

end

Syntax Directed Translation Schemes (SDT’S)

 It is a CFG with program fragments embedded with the production body.

 The program fragments are called semantic actions

SDD for infix to postfix translation

GRAMMAR SEMANTIC ACTIONS

15
Bottom up evaluation of S-attributed definition
 Syntax-directed definitions with only synthesized attributes(S- attribute) can be evaluated
by a bottom up parser (BUP) as input is parsed
 In this approach, the parser will keep the values of synthesized attributes associated with
the grammar symbol on its stack.
 The stack is implemented as a pair of state and value.
 When a reduction is made ,the values of the synthesized attributes are computed from the
attribute appearing on the stack for the grammar symbols
 implementation is by using an LR parser (e.g. YACC)

Consider again the syntax-directed definition of the desk calculator.

16
Example :- SDD and code fragment using S attributed definition for the input 3*5+4n is as
follows:

17
Top Down Translation

Do left
recursion for this grammar

18
Type Checking

 The main aim of the compiler is to translate the source program to a form that can be
executed on a target machine.
 For this purpose the compiler need to
1. Check that the source program follows the syntax and semantics of the concerned
language.
2. Check the flow of data between variables.

What is type?

 Type is a notion or a rule that varies from language to language.

 So type checking is done to ensure that whether the source program follows the syntactic
and semantic rule of that language.
 Type checking can be of two types:
1. Static checking (Done at compile time)
2. Dynamic checking. (Done during run time)

Static checking(Semantic check)

 Its helps in detecting and reporting program errors.

 Some of the static check includes:
1. Type checks
2. Flow of control checks
3. Uniqueness check – object must be defined exactly one for some scenarios
4. Name related check-Same name must appear two or more times
Eg: function call and definition.

19
Type expressions

 Type of a language construct will be denoted by type expression

 Type expression can be:
1. Basic type
2. Type expression formed by applying an operator called a type constructor to other
type expression.

20
21
Specification of simple type checker

A simple language

22
 This grammar will generate programs represented by the nonterminal P, consisting of
a sequence of declaration D followed by a single expression E.

Type checking for expressions

Type checking for statements

 Statements do not have values, therefore a special type void can be assigned to them.
 If an error occurs within a statement, the type assigned to the statement is type_error.

S id= E { S.type=(if id.type=E.type then void) else type error}

S if E then S1 { S.type=(if E.type=Boolean then S1.type) else type error}

23
SWhile E do S1 { S.type=(if E.type=Boolean then S1.type) else type error}

SS1;S2 { S.type=(if S1.type=void and S2.type=void then void) else

type error}

Type checking for functions

EE1(E2) {E.type=(if E2.type=S and E1.type=st then t) else type

error}

TT1T2 {T.type=T1.typeT2.type}

**********

24
UNIT-3 INTERMEDIATE CODE GENERATION

INTRODUCTION

The front end translates a source program into an intermediate representation from which
the back end generates target code.

Benefits of using a machine-independent intermediate form are:

1. Retargeting is facilitated. That is, a compiler for a different machine can be created by
attaching a back end for the new machine to an existing front end.

2. A machine-independent code optimizer can be applied to the intermediate representation.

Position of intermediate code generator

parser static intermediate intermediate code

checker code generator generator
code

INTERMEDIATE LANGUAGES

Three ways of intermediate representation:

 Syntax tree

 Postfix notation

 Three address code

The semantic rules for generating three-address code from common programming language
constructs are similar to those for constructing syntax trees or for generating postfix notation.

Graphical Representations:

Syntax tree:

A syntax tree depicts the natural hierarchical structure of a source program. A dag
(Directed Acyclic Graph) gives the same information but in a more compact way because
common subexpressions are identified. A syntax tree and dag for the assignment statement a : =
b * - c + b * - c are as follows:
assign assign

a + a +

* * *

b uminus b uminus b uminus

c c c

(a) Syntax tree (b) Dag

Postfix notation:

Postfix notation is a linearized representation of a syntax tree; it is a list of the nodes of

the tree in which a node appears immediately after its children. The postfix notation for the
syntax tree given above is

a b c uminus * b c uminus * + assign

Syntax-directed definition:

Syntax trees for assignment statements are produced by the syntax-directed definition.
Non-terminal S generates an assignment statement. The two binary operators + and * are
examples of the full operator set in a typical language. Operator associativities and precedences
are the usual ones, even though they have not been put into the grammar. This definition
constructs the tree from the input a : = b * - c + b* - c.

PRODUCTION SEMANTIC RULE

S  id : = E S.nptr : = mknode(‘assign’,mkleaf(id, id.place), E.nptr)

E  E1 + E2 E.nptr : = mknode(‘+’, E1.nptr, E2.nptr )

E  E1 * E2 E.nptr : = mknode(‘*’, E 1.nptr, E2.nptr )

E  - E1 E.nptr : = mknode(‘uminus’, E1.nptr)

E  ( E1 ) E.nptr : = E1.nptr

E  id E.nptr : = mkleaf( id, id.place )

Syntax-directed definition to produce syntax trees for assignment statements

The token id has an attribute place that points to the symbol-table entry for the identifier.
A symbol-table entry can be found from an attribute id.name, representing the lexeme associated
with that occurrence of id. If the lexical analyzer holds all lexemes in a single array of
characters, then attribute name might be the index of the first character of the lexeme.

Two representations of the syntax tree are as follows. In (a) each node is represented as a
record with a field for its operator and additional fields for pointers to its children. In (b), nodes
are allocated from an array of records and the index or position of the node serves as the pointer
to the node. All the nodes in the syntax tree can be visited by following pointers, starting from
the root at position 10.

Two representations of the syntax tree

aaaaaaaaaaaaa
assign 0 id b

1 id c
id a
2 uminus
2 1

3 * 0 2
+
4 id b

5 id c
* *
6 uminus 5
id b id b
7 * 4 6

uminus uminus 8 + 3 7

9 id a
id c id c
10 assign 9 8

(a) (b)

Three-Address Code:

Three-address code is a sequence of statements of the general form

x : = y op z

where x, y and z are names, constants, or compiler-generated temporaries; op stands for any
operator, such as a fixed- or floating-point arithmetic operator, or a logical operator on boolean-
valued data. Thus a source language expression like x+ y*z might be translated into a sequence

t1 : = y * z
t2 : = x + t 1

where t1 and t2 are compiler-generated temporary names.

Advantages of three-address code:

 The unraveling of complicated arithmetic expressions and of nested flow-of-control

statements makes three-address code desirable for target code generation and
optimization.

 The use of names for the intermediate values computed by a program allows three-
address code to be easily rearranged – unlike postfix notation.

Three-address code is a linearized representation of a syntax tree or a dag in which

explicit names correspond to the interior nodes of the graph. The syntax tree and dag are
represented by the three-address code sequences. Variable names can appear directly in three-
address statements.

Three-address code corresponding to the syntax tree and dag given above

t1 : = - c t 1 : = -c

t 2 : = b * t1 t2 : = b * t1

t3 : = - c t 5 : = t2 + t2

t 4 : = b * t3 a : = t5

t5 : = t2 + t4

a : = t5

(a) Code for the syntax tree (b) Code for the dag

The reason for the term “three-address code” is that each statement usually contains three
addresses, two for the operands and one for the result.

Types of Three-Address Statements:

The common three-address statements are:

1. Assignment statements of the form x : = y op z, where op is a binary arithmetic or logical

operation.

2. Assignment instructions of the form x : = op y, where op is a unary operation. Essential unary

operations include unary minus, logical negation, shift operators, and conversion operators
that, for example, convert a fixed-point number to a floating-point number.

3. Copy statements of the form x : = y where the value of y is assigned to x.

4. The unconditional jump goto L. The three-address statement with label L is the next to be
executed.

5. Conditional jumps such as if x relop y goto L. This instruction applies a relational operator (
<, =, >=, etc. ) to x and y, and executes the statement with label L next if x stands in relation
relop to y. If not, the three-address statement following if x relop y goto L is executed next,
as in the usual sequence.

6. param x and call p, n for procedure calls and return y, where y representing a returned value
is optional. For example,
param x1
param x2
...
param xn
call p,n
generated as part of a call of the procedure p(x 1, x2, …. ,xn ).

7. Indexed assignments of the form x : = y[i] and x[i] : = y.

8. Address and pointer assignments of the form x : = &y , x : = y, and x : = y.

Syntax-Directed Translation into Three-Address Code:

When three-address code is generated, temporary names are made up for the interior
nodes of a syntax tree. For example, id : = E consists of code to evaluate E into some temporary
t, followed by the assignment id.place : = t.

Given input a : = b * - c + b * - c, the three-address code is as shown above. The

synthesized attribute S.code represents the three-address code for the assignment S.
The nonterminal E has two attributes :
1. E.place, the name that will hold the value of E , and
2. E.code, the sequence of three-address statements evaluating E.

Syntax-directed definition to produce three-address code for assignments

PRODUCTION SEMANTIC RULES

S  id : = E S.code : = E.code || gen(id.place ‘:=’ E.place)

E  E1 + E2 E.place := newtemp;
E.code := E1.code || E2.code || gen(E.place ‘:=’ E1.place ‘+’ E2.place)

E  E1 * E2 E.place := newtemp;
E.code := E1.code || E2.code || gen(E.place ‘:=’ E 1.place ‘*’ E2.place)

E  - E1 E.place := newtemp;
E.code := E1.code || gen(E.place ‘:=’ ‘uminus’ E1.place)

E  ( E1 ) E.place : = E1.place;
E.code : = E1.code

E  id E.place : = id.place;
E.code : = ‘ ‘
Semantic rules generating code for a while statement

S.begin:

E.code

if E.place = 0 goto S.after

S1.code

goto S.begin

S.after: ...

PRODUCTION SEMANTIC RULES

S  while E do S1 S.begin := newlabel;

S.after := newlabel;
S.code := gen(S.begin ‘:’) ||
E.code ||
gen ( ‘if’ E.place ‘=’ ‘0’ ‘goto’ S.after)||
S1.code ||
gen ( ‘goto’ S.begin) ||
gen ( S.after ‘:’)

 The function newtemp returns a sequence of distinct names t 1,t2,….. in response to

successive calls.
 Notation gen(x ‘:=’ y ‘+’ z) is used to represent three-address statement x := y + z.
Expressions appearing instead of variables like x, y and z are evaluated when passed to
gen, and quoted operators or operand, like ‘+’ are taken literally.
 Flow-of–control statements can be added to the language of assignments. The code for S
 while E do S1 is generated using new attributes S.begin and S.after to mark the first
statement in the code for E and the statement following the code for S, respectively.
 The function newlabel returns a new label every time it is called.
 We assume that a non-zero expression represents true; that is when the value of E
becomes zero, control leaves the while statement.

Implementation of Three-Address Statements:

A three-address statement is an abstract form of intermediate code. In a compiler,

these statements can be implemented as records with fields for the operator and the operands.
Three such representations are:
 Quadruples

 Triples

 Indirect triples

Quadruples:

 A quadruple is a record structure with four fields, which are, op, arg1, arg2 and result.

 The op field contains an internal code for the operator. The three-address statement x : =
y op z is represented by placing y in arg1, z in arg2 and x in result.

 The contents of fields arg1, arg2 and result are normally pointers to the symbol-table
entries for the names represented by these fields. If so, temporary names must be entered
into the symbol table as they are created.

Triples:

 To avoid entering temporary names into the symbol table, we might refer to a temporary
value by the position of the statement that computes it.

 If we do so, three-address statements can be represented by records with only three fields:
op, arg1 and arg2.

 The fields arg1 and arg2, for the arguments of op, are either pointers to the symbol table
or pointers into the triple structure ( for temporary values ).

 Since three fields are used, this intermediate code format is known as triples.

op arg1 arg2 result op arg1 arg2

(0) uminus c t1 (0) uminus c

(1) * b t1 t2 (1) * b (0)

(2) uminus c t3 (2) uminus c

(3) * b t3 t4 (3) * b (2)

(4) + t2 t4 t5 (4) + (1) (3)

(5) := t3 a (5) assign a (4)

(a) Quadruples (b) Triples

Quadruple and triple representation of three-address statements given above

A ternary operation like x[i] : = y requires two entries in the triple structure as shown as below
while x : = y[i] is naturally represented as two operations.

op arg1 arg2 op arg1 arg2

(0) []= x i (0) =[] y i

(1) assign (0) y (1) assign x (0)

(a) x[i] : = y (b) x : = y[i]

Indirect Triples:

 Another implementation of three-address code is that of listing pointers to triples, rather

than listing the triples themselves. This implementation is called indirect triples.

 For example, let us use an array statement to list pointers to triples in the desired order.
Then the triples shown above might be represented as follows:

statement op arg1 arg2

(0) (14) (14) uminus c

(1) (15) (15) * b (14)
(2) (16) (16) uminus c
(3) (17) (17) * b (16)
(4) (18) (18) + (15) (17)
(5) (19) (19) assign a (18)

Indirect triples representation of three-address statements

DECLARATIONS

As the sequence of declarations in a procedure or block is examined, we can lay out

storage for names local to the procedure. For each local name, we create a symbol-table entry
with information like the type and the relative address of the storage for the name. The relative
address consists of an offset from the base of the static data area or the field for local data in an
activation record.
Declarations in a Procedure:
The syntax of languages such as C, Pascal and Fortran, allows all the declarations in a
single procedure to be processed as a group. In this case, a global variable, say offset, can keep
track of the next available relative address.

In the translation scheme shown below:

 Nonterminal P generates a sequence of declarations of the form id : T.

 Before the first declaration is considered, offset is set to 0. As each new name is seen ,
that name is entered in the symbol table with offset equal to the current value of offset,
and offset is incremented by the width of the data object denoted by that name.

 The procedure enter( name, type, offset ) creates a symbol-table entry for name, gives its
type type and relative address offset in its data area.

 Attribute type represents a type expression constructed from the basic types integer and
real by applying the type constructors pointer and array. If type expressions are
represented by graphs, then attribute type might be a pointer to the node representing a
type expression.

 The width of an array is obtained by multiplying the width of each element by the
number of elements in the array. The width of each pointer is assumed to be 4.

Computing the types and relative addresses of declared names

P D { offset : = 0 }

DD;D

D  id : T { enter(id.name, T.type, offset);

offset : = offset + T.width }

T  integer { T.type : = integer;

T.width : = 4 }

T  real { T.type : = real;

T.width : = 8 }

T  array [ num ] of T1 { T.type : = array(num.val, T 1.type);

T.width : = num.val X T1.width }

T  ↑ T1 { T.type : = pointer ( T1.type);

T.width : = 4 }
Keeping Track of Scope Information:

When a nested procedure is seen, processing of declarations in the enclosing procedure is

temporarily suspended. This approach will be illustrated by adding semantic rules to the
following language:

PD

D  D ; D | id : T | proc id ; D ; S

One possible implementation of a symbol table is a linked list of entries for names.

A new symbol table is created when a procedure declaration D  proc id D1;S is seen,
and entries for the declarations in D 1 are created in the new table. The new table points back to
the symbol table of the enclosing procedure; the name represented by id itself is local to the
enclosing procedure. The only change from the treatment of variable declarations is that the
procedure enter is told which symbol table to make an entry in.

For example, consider the symbol tables for procedures readarray, exchange, and
quicksort pointing back to that for the containing procedure sort, consisting of the entire
program. Since partition is declared within quicksort, its table points to that of quicksort.

Symbol tables for nested procedures

sort
nil header
a
x
readarray to readarray
exchange to exchange
quicksort

readarray exchange quicksort

header header header
i k
v
partition

partition
header
i
j
The semantic rules are defined in terms of the following operations:

1. mktable(previous) creates a new symbol table and returns a pointer to the new table. The
argument previous points to a previously created symbol table, presumably that for the
enclosing procedure.

2. enter(table, name, type, offset) creates a new entry for name name in the symbol table pointed
to by table. Again, enter places type type and relative address offset in fields within the entry.

3. addwidth(table, width) records the cumulative width of all the entries in table in the header
associated with this symbol table.

4. enterproc(table, name, newtable) creates a new entry for procedure name in the symbol table
pointed to by table. The argument newtable points to the symbol table for this procedure
name.

Syntax directed translation scheme for nested procedures

PMD { addwidth ( top( tblptr) , top (offset));

pop (tblptr); pop (offset) }

Mɛ { t : = mktable (nil);

push (t,tblptr); push (0,offset) }

D  D1 ; D2

D  proc id ; N D1 ; S { t : = top (tblptr);

addwidth ( t, top (offset));
pop (tblptr); pop (offset);
enterproc (top (tblptr), id.name, t) }

D  id : T { enter (top (tblptr), id.name, T.type, top (offset));

top (offset) := top (offset) + T.width }

Nɛ { t := mktable (top (tblptr));

push (t, tblptr); push (0,offset) }

 The stack tblptr is used to contain pointers to the tables for sort, quicksort, and partition
when the declarations in partition are considered.

 The top element of stack offset is the next available relative address for a local of the
current procedure.

 All semantic actions in the subtrees for B and C in

A  BC {actionA}

are done before actionA at the end of the production occurs. Hence, the action associated
with the marker M is the first to be done.
 The action for nonterminal M initializes stack tblptr with a symbol table for the
outermost scope, created by operation mktable(nil). The action also pushes relative
address 0 onto stack offset.

 Similarly, the nonterminal N uses the operation mktable(top(tblptr)) to create a new

symbol table. The argument top(tblptr) gives the enclosing scope for the new table.

 For each variable declaration id: T, an entry is created for id in the current symbol table.
The top of stack offset is incremented by T.width.

 When the action on the right side of D  proc id; ND1; S occurs, the width of all
declarations generated by D1 is on the top of stack offset; it is recorded using addwidth.
Stacks tblptr and offset are then popped.
At this point, the name of the enclosed procedure is entered into the symbol table of its
enclosing procedure.

ASSIGNMENT STATEMENTS

Suppose that the context in which an assignment appears is given by the following grammar.

PMD

Mɛ

D  D ; D | id : T | proc id ; N D ; S

Nɛ

Nonterminal P becomes the new start symbol when these productions are added to those in the
translation scheme shown below.

Translation scheme to produce three-address code for assignments

S  id : = E { p : = lookup ( id.name);
if p ≠ nil then
emit( p ‘ : =’ E.place)
else error }

E  E1 + E2 { E.place : = newtemp;
emit( E.place ‘: =’ E1.place ‘ + ‘ E2.place ) }

E  E1 * E2 { E.place : = newtemp;
emit( E.place ‘: =’ E1.place ‘ * ‘ E2.place ) }

E  - E1 { E.place : = newtemp;
emit ( E.place ‘: =’ ‘uminus’ E1.place ) }

E  ( E1 ) { E.place : = E1.place }
E  id { p : = lookup ( id.name);

if p ≠ nil then
E.place : = p
else error }

Reusing Temporary Names

 The temporaries used to hold intermediate values in expression calculations tend to

clutter up the symbol table, and space has to be allocated to hold their values.

 Temporaries can be reused by changing newtemp. The code generated by the rules for E
 E1 + E2 has the general form:

evaluate E1 into t1
evaluate E2 into t2
t : = t 1 + t2

 The lifetimes of these temporaries are nested like matching pairs of balanced parentheses.

 Keep a count c , initialized to zero. Whenever a temporary name is used as an operand,

decrement c by 1. Whenever a new temporary name is generated, use $c and increase c
by 1.

 For example, consider the assignment x := a * b + c * d – e * f

Three-address code with stack temporaries

statement value of c

0
$0 := a * b 1
$1 := c * d 2
$0 := $0 + $1 1
$1 := e * f 2
$0 := $0 - $1 1
x := $0 0

Addressing Array Elements:

Elements of an array can be accessed quickly if the elements are stored in a block of
consecutive locations. If the width of each array element is w, then the ith element of array A
begins in location

base + ( i – low ) x w

where low is the lower bound on the subscript and base is the relative address of the storage
allocated for the array. That is, base is the relative address of A[low].
The expression can be partially evaluated at compile time if it is rewritten as

i x w + ( base – low x w)

The subexpression c = base – low x w can be evaluated when the declaration of the array is seen.
We assume that c is saved in the symbol table entry for A , so the relative address of A[i] is
obtained by simply adding i x w to c.

Address calculation of multi-dimensional arrays:

A two-dimensional array is stored in of the two forms :

 Row-major (row-by-row)

 Column-major (column-by-column)

Layouts for a 2 x 3 array

A[ 1,1 ] A [ 1,1 ]
first column
first row A[ 1,2 ] A [ 2,1 ]
A[ 1,3 ] A [ 1,2 ]
A[ 2,1 ] A [ 2,2 ] second column

second row A[ 2,2 ] A [ 1,3 ]

A[ 2,3 ] A [ 2,3 ] third column

(a) ROW-MAJOR (b) COLUMN-MAJOR

In the case of row-major form, the relative address of A[ i1 , i2] can be calculated by the formula

base + ((i1 – low1) x n2 + i2 – low2) x w

where, low1 and low2 are the lower bounds on the values of i1 and i2 and n2 is the number of
values that i2 can take. That is, if high2 is the upper bound on the value of i2, then n2 = high2 –
low2 + 1.

Assuming that i1 and i2 are the only values that are known at compile time, we can rewrite the
above expression as

(( i1 x n2 ) + i2 ) x w + ( base – (( low1 x n2 ) + low2 ) x w)

Generalized formula:

The expression generalizes to the following expression for the relative address of A[i 1,i2,…,ik]

(( . . . (( i1n2 + i2 ) n3 + i3) . . . ) nk + ik ) x w + base – (( . . .((low1n2 + low2)n3 + low3) . . .)

nk + lowk) x w

for all j, nj = highj – lowj + 1

The Translation Scheme for Addressing Array Elements :

Semantic actions will be added to the grammar :

(1) S  L:=E
(2) E  E+E
(3) E  (E)
(4) E  L
(5) L  Elist ]
(6) L  id
(7) Elist  Elist , E
(8) Elist  id [ E

We generate a normal assignment if L is a simple name, and an indexed assignment into the
location denoted by L otherwise :

(1) S  L : = E { if L.offset = null then / * L is a simple id */

emit ( L.place ‘: =’ E.place ) ;
else
emit ( L.place ‘ [‘ L.offset ‘ ]’ ‘: =’ E.place) }

(2) E  E1 + E2 { E.place : = newtemp;

emit ( E.place ‘: =’ E1.place ‘ +’ E2.place ) }

(3) E  ( E1 ) { E.place : = E1.place }

When an array reference L is reduced to E , we want the r-value of L. Therefore we use indexing
to obtain the contents of the location L.place [ L.offset ] :

(4) E  L { if L.offset = null then /* L is a simple id* /

E.place : = L.place
else begin
E.place : = newtemp;
emit ( E.place ‘: =’ L.place ‘ [‘ L.offset ‘]’)
end }

(5) L  Elist ] { L.place : = newtemp;

L.offset : = newtemp;
emit (L.place ‘: =’ c( Elist.array ));
emit (L.offset ‘: =’ Elist.place ‘*’ width (Elist.array)) }

(6) L  id { L.place := id.place;

L.offset := null }

(7) Elist  Elist1 , E { t := newtemp;

m : = Elist1.ndim + 1;
emit ( t ‘: =’ Elist1.place ‘*’ limit (Elist1.array,m));
emit ( t ‘: =’ t ‘+’ E.place);
Elist.array : = Elist1.array;
Elist.place : = t;
Elist.ndim : = m }

(8) Elist  id [ E { Elist.array : = id.place;

Elist.place : = E.place;
Elist.ndim : = 1 }

Type conversion within Assignments :

Consider the grammar for assignment statements as above, but suppose there are two
types – real and integer , with integers converted to reals when necessary. We have another
attribute E.type, whose value is either real or integer. The semantic rule for E.type associated
with the production E  E + E is :

EE+E { E.type : =
if E1.type = integer and
E2.type = integer then integer
else real }

The entire semantic rule for E  E + E and most of the other productions must be
modified to generate, when necessary, three-address statements of the form x : = inttoreal y,
whose effect is to convert integer y to a real of equal value, called x.

Semantic action for E  E1 + E2

E.place := newtemp;
if E1.type = integer and E2.type = integer then begin
emit( E.place ‘: =’ E1.place ‘int +’ E2.place);
E.type : = integer
end
else if E1.type = real and E2.type = real then begin
emit( E.place ‘: =’ E1.place ‘real +’ E2.place);
E.type : = real
end
else if E1.type = integer and E2.type = real then begin
u : = newtemp;
emit( u ‘: =’ ‘inttoreal’ E1.place);
emit( E.place ‘: =’ u ‘ real +’ E2.place);
E.type : = real
end
else if E1.type = real and E2.type =integer then begin
u : = newtemp;
emit( u ‘: =’ ‘inttoreal’ E2.place);
emit( E.place ‘: =’ E1.place ‘ real +’ u);
E.type : = real
end
else
E.type : = type_error;
For example, for the input x : = y + i * j
assuming x and y have type real, and i and j have type integer, the output would look like

t1 : = i int* j
t3 : = inttoreal t1
t2 : = y real+ t3
x : = t2

BOOLEAN EXPRESSIONS

Boolean expressions have two primary purposes. They are used to compute logical
values, but more often they are used as conditional expressions in statements that alter the flow
of control, such as if-then-else, or while-do statements.

Boolean expressions are composed of the boolean operators ( and, or, and not ) applied
to elements that are boolean variables or relational expressions. Relational expressions are of the
form E1 relop E2, where E1 and E2 are arithmetic expressions.

Here we consider boolean expressions generated by the following grammar :

E  E or E | E and E | not E | ( E ) | id relop id | true | false

Methods of Translating Boolean Expressions:

There are two principal methods of representing the value of a boolean expression. They are :

 To encode true and false numerically and to evaluate a boolean expression analogously
to an arithmetic expression. Often, 1 is used to denote true and 0 to denote false.

 To implement boolean expressions by flow of control, that is, representing the value of a
boolean expression by a position reached in a program. This method is particularly
convenient in implementing the boolean expressions in flow-of-control statements, such
as the if-then and while-do statements.

Numerical Representation

Here, 1 denotes true and 0 denotes false. Expressions will be evaluated completely from
left to right, in a manner similar to arithmetic expressions.

For example :

 The translation for

a or b and not c
is the three-address sequence
t1 : = not c
t2 : = b and t1
t3 : = a or t2

 A relational expression such as a < b is equivalent to the conditional statement

if a < b then 1 else 0
which can be translated into the three-address code sequence (again, we arbitrarily start
statement numbers at 100) :

100 : if a < b goto 103

101 : t:=0
102 : goto 104
103 : t:=1
104 :

Translation scheme using a numerical representation for booleans

E  E1 or E2 { E.place : = newtemp;
emit( E.place ‘: =’ E1.place ‘or’ E2.place ) }
E  E1 and E2 { E.place : = newtemp;
emit( E.place ‘: =’ E1.place ‘and’ E2.place ) }
E  not E1 { E.place : = newtemp;
emit( E.place ‘: =’ ‘not’ E1.place ) }
E  ( E1 ) { E.place : = E1.place }
E  id1 relop id2 { E.place : = newtemp;
emit( ‘if’ id1.place relop.op id2.place ‘goto’ nextstat + 3);
emit( E.place ‘: =’ ‘0’ );
emit(‘goto’ nextstat +2);
emit( E.place ‘: =’ ‘1’) }
E  true { E.place : = newtemp;
emit( E.place ‘: =’ ‘1’) }
E false { E.place : = newtemp;
emit( E.place ‘: =’ ‘0’) }

Short-Circuit Code:

We can also translate a boolean expression into three-address code without generating
code for any of the boolean operators and without having the code necessarily evaluate the entire
expression. This style of evaluation is sometimes called “short-circuit” or “jumping” code. It is
possible to evaluate boolean expressions without generating code for the boolean operators and,
or, and not if we represent the value of an expression by a position in the code sequence.

Translation of a < b or c < d and e < f

100 : if a < b goto 103 107 : t2 : = 1

101 : t1 : = 0 108 : if e < f goto 111

102 : goto 104 109 : t3 : = 0

103 : t1 : = 1 110 : goto 112

104 : if c < d goto 107 111 : t3 : = 1

105 : t2 : = 0 112 : t4 : = t2 and t3

106 : goto 108 113 : t5 : = t1 or t4

Flow-of-Control Statements

We now consider the translation of boolean expressions into three-address code in the
context of if-then, if-then-else, and while-do statements such as those generated by the following
grammar:

S  if E then S1
| if E then S1 else S2
| while E do S1

In each of these productions, E is the Boolean expression to be translated. In the translation, we

assume that a three-address statement can be symbolically labeled, and that the function
newlabel returns a new symbolic label each time it is called.

 E.true is the label to which control flows if E is true, and E.false is the label to which
control flows if E is false.

 The semantic rules for translating a flow-of-control statement S allow control to flow
from the translation S.code to the three-address instruction immediately following
S.code.

 S.next is a label that is attached to the first three-address instruction to be executed after
the code for S.

Code for if-then , if-then-else, and while-do statements

to E.true
E.code
to E.false
E.code to E.true E.true: S1.code

E.true : to E.false
S1.code goto S.next
E.false:
S2.code
E.false : ...

S.next: ...

(a) if-then (b) if-then-else

S.begin: E.code to E.true

to E.false
E.true: S1.code

goto S.begin
E.false: ...

PRODUCTION SEMANTIC RULES

S  if E then S1 E.true : = newlabel;

E.false : = S.next;
S1.next : = S.next;
S.code : = E.code || gen(E.true ‘:’) || S 1.code

S  if E then S1 else S2 E.true : = newlabel;

E.false : = newlabel;
S1.next : = S.next;
S2.next : = S.next;
S.code : = E.code || gen(E.true ‘:’) || S 1.code ||
gen(‘goto’ S.next) ||
gen( E.false ‘:’) || S2.code

S  while E do S1 S.begin : = newlabel;

E.true : = newlabel;
E.false : = S.next;
S1.next : = S.begin;
S.code : = gen(S.begin ‘:’)|| E.code ||
gen(E.true ‘:’) || S1.code ||
gen(‘goto’ S.begin)

Control-Flow Translation of Boolean Expressions:

Syntax-directed definition to produce three-address code for booleans

PRODUCTION SEMANTIC RULES

E  E1 or E2 E1.true : = E.true;
E1.false : = newlabel;
E2.true : = E.true;
E2.false : = E.false;
E.code : = E1.code || gen(E1.false ‘:’) || E2.code

E  E1 and E2 E.true : = newlabel;

E1.false : = E.false;
E2.true : = E.true;
E2.false : = E.false;
E.code : = E1.code || gen(E1.true ‘:’) || E2.code

E  not E1 E1.true : = E.false;

E1.false : = E.true;
E.code : = E1.code

E  ( E1 ) E1.true : = E.true;
E1.false : = E.false;
E.code : = E1.code

E  id1 relop id2 E.code : = gen(‘if’ id1.place relop.op id2.place

‘goto’ E.true) || gen(‘goto’ E.false)

E  true E.code : = gen(‘goto’ E.true)

E  false E.code : = gen(‘goto’ E.false)

CASE STATEMENTS

The “switch” or “case” statement is available in a variety of languages. The switch-statement

syntax is as shown below :
Switch-statement syntax

switch expression
begin
case value : statement
case value : statement
...
case value : statement
default : statement
end

There is a selector expression, which is to be evaluated, followed by n constant values

that the expression might take, including a default “value” which always matches the expression
if no other value does. The intended translation of a switch is code to:

1. Evaluate the expression.

2. Find which value in the list of cases is the same as the value of the expression.
3. Execute the statement associated with the value found.

Step (2) can be implemented in one of several ways :

 By a sequence of conditional goto statements, if the number of cases is small.

 By creating a table of pairs, with each pair consisting of a value and a label for the code
of the corresponding statement. Compiler generates a loop to compare the value of the
expression with each value in the table. If no match is found, the default (last) entry is
sure to match.
 If the number of cases s large, it is efficient to construct a hash table.
 There is a common special case in which an efficient implementation of the n-way branch
exists. If the values all lie in some small range, say i min to imax, and the number of
different values is a reasonable fraction of i max - imin, then we can construct an array of
labels, with the label of the statement for value j in the entry of the table with offset j -
imin and the label for the default in entries not filled otherwise. To perform switch,
evaluate the expression to obtain the value of j , check the value is within range and
transfer to the table entry at offset j-imin .

Syntax-Directed Translation of Case Statements:

Consider the following switch statement:

switch E
begin
case V1 : S1
case V2 : S2
...
case Vn-1 : Sn-1
default : Sn
end

This case statement is translated into intermediate code that has the following form :

Translation of a case statement

code to evaluate E into t

goto test
L1 : code for S1
goto next
L2 : code for S2
goto next
. . .
Ln-1 : code for Sn-1
goto next
Ln : code for Sn
goto next
test : if t = V1 goto L1
if t = V2 goto L2
. . .
if t = Vn-1 goto Ln-1
goto Ln
next :

To translate into above form :

 When keyword switch is seen, two new labels test and next, and a new temporary t are
generated.

 As expression E is parsed, the code to evaluate E into t is generated. After processing E ,

the jump goto test is generated.

 As each case keyword occurs, a new label Li is created and entered into the symbol table.
A pointer to this symbol-table entry and the value Vi of case constant are placed on a
stack (used only to store cases).
 Each statement case Vi : Si is processed by emitting the newly created label Li, followed
by the code for Si , followed by the jump goto next.

 Then when the keyword end terminating the body of the switch is found, the code can be
generated for the n-way branch. Reading the pointer-value pairs on the case stack from
the bottom to the top, we can generate a sequence of three-address statements of the form

case V1 L1
case V2 L2
...
case Vn-1 Ln-1
case t Ln
label next

where t is the name holding the value of the selector expression E, and Ln is the label for
the default statement.

BACKPATCHING

The easiest way to implement the syntax-directed definitions for boolean expressions is
to use two passes. First, construct a syntax tree for the input, and then walk the tree in depth-first
order, computing the translations. The main problem with generating code for boolean
expressions and flow-of-control statements in a single pass is that during one single pass we may
not know the labels that control must go to at the time the jump statements are generated. Hence,
a series of branching statements with the targets of the jumps left unspecified is generated. Each
statement will be put on a list of goto statements whose labels will be filled in when the proper
label can be determined. We call this subsequent filling in of labels backpatching.

To manipulate lists of labels, we use three functions :

1. makelist(i) creates a new list containing only i, an index into the array of quadruples;
makelist returns a pointer to the list it has made.
2. merge(p1,p2) concatenates the lists pointed to by p1 and p2, and returns a pointer to the
concatenated list.
3. backpatch(p,i) inserts i as the target label for each of the statements on the list pointed to
by p.

Boolean Expressions:

We now construct a translation scheme suitable for producing quadruples for boolean
expressions during bottom-up parsing. The grammar we use is the following:

(1) E  E1 or M E2
(2) | E1 and M E2
(3) | not E1
(4) | ( E1 )
(5) | id1 relop id2
(6) | true
(7) | false
(8) M  ɛ
Synthesized attributes truelist and falselist of nonterminal E are used to generate jumping code
for boolean expressions. Incomplete jumps with unfilled labels are placed on lists pointed to by
E.truelist and E.falselist.

Consider production E  E1 and M E2. If E1 is false, then E is also false, so the statements on
E1.falselist become part of E.falselist. If E1 is true, then we must next test E2, so the target for the
statements E1.truelist must be the beginning of the code generated for E2. This target is obtained
using marker nonterminal M.

Attribute M.quad records the number of the first statement of E2.code. With the production M 
ɛ we associate the semantic action

{ M.quad : = nextquad }

The variable nextquad holds the index of the next quadruple to follow. This value will be
backpatched onto the E1.truelist when we have seen the remainder of the production E  E1 and
M E2. The translation scheme is as follows:

(1) E  E1 or M E2 { backpatch ( E1.falselist, M.quad);

E.truelist : = merge( E1.truelist, E2.truelist);
E.falselist : = E2.falselist }

(2) E  E1 and M E2 { backpatch ( E1.truelist, M.quad);

E.truelist : = E2.truelist;
E.falselist : = merge(E1.falselist, E2.falselist) }

(3) E  not E1 { E.truelist : = E1.falselist;

E.falselist : = E1.truelist; }

(4) E  ( E1 ) { E.truelist : = E1.truelist;

E.falselist : = E1.falselist; }

(5) E  id1 relop id2 { E.truelist : = makelist (nextquad);

E.falselist : = makelist(nextquad + 1);
emit(‘if’ id1.place relop.op id2.place ‘goto_’)
emit(‘goto_’) }

(6) E  true { E.truelist : = makelist(nextquad);

emit(‘goto_’) }

(7) E  false { E.falselist : = makelist(nextquad);

emit(‘goto_’) }

(8) M  ɛ { M.quad : = nextquad }

Flow-of-Control Statements:

A translation scheme is developed for statements generated by the following grammar :

Here S denotes a statement, L a statement list, A an assignment statement, and E a boolean

expression. We make the tacit assumption that the code that follows a given statement in
execution also follows it physically in the quadruple array. Else, an explicit jump must be
provided.

Scheme to implement the Translation:

The nonterminal E has two attributes E.truelist and E.falselist. L and S also need a list of
unfilled quadruples that must eventually be completed by backpatching. These lists are pointed
to by the attributes L..nextlist and S.nextlist. S.nextlist is a pointer to a list of all conditional and
unconditional jumps to the quadruple following the statement S in execution order, and L.nextlist
is defined similarly.

The semantic rules for the revised grammar are as follows:

(1) S  if E then M1 S1 N else M2 S2

{ backpatch (E.truelist, M1.quad);
backpatch (E.falselist, M2.quad);
S.nextlist : = merge (S1.nextlist, merge (N.nextlist, S2.nextlist)) }

We backpatch the jumps when E is true to the quadruple M1.quad, which is the beginning of the
code for S1. Similarly, we backpatch jumps when E is false to go to the beginning of the code for
S2. The list S.nextlist includes all jumps out of S 1 and S2, as well as the jump generated by N.

(2) Nɛ { N.nextlist : = makelist( nextquad );

emit(‘goto _’) }

(3) Mɛ { M.quad : = nextquad }

(4) S  if E then M S1 { backpatch( E.truelist, M.quad);

S.nextlist : = merge( E.falselist, S1.nextlist) }

(5) S  while M1 E do M2 S1 { backpatch( S1.nextlist, M1.quad);

backpatch( E.truelist, M2.quad);
S.nextlist : = E.falselist
emit( ‘goto’ M1.quad ) }

(6) S  begin L end { S.nextlist : = L.nextlist }

(7) SA { S.nextlist : = nil }

The assignment S.nextlist : = nil initializes S.nextlist to an empty list.

(8) L  L1 ; M S { backpatch( L1.nextlist, M.quad);

L.nextlist : = S.nextlist }

The statement following L1 in order of execution is the beginning of S. Thus the L1.nextlist list is
backpatched to the beginning of the code for S, which is given by M.quad.

(9) LS { L.nextlist : = S.nextlist }

PROCEDURE CALLS

The procedure is such an important and frequently used programming construct that it is
imperative for a compiler to generate good code for procedure calls and returns. The run-time
routines that handle procedure argument passing, calls and returns are part of the run-time
support package.

Let us consider a grammar for a simple procedure call statement

(1) S  call id ( Elist )

(2) Elist  Elist , E
(3) Elist  E

Calling Sequences:

The translation for a call includes a calling sequence, a sequence of actions taken on entry
to and exit from each procedure. The falling are the actions that take place in a calling sequence :

 When a procedure call occurs, space must be allocated for the activation record of the
called procedure.

 The arguments of the called procedure must be evaluated and made available to the called
procedure in a known place.

 Environment pointers must be established to enable the called procedure to access data in
enclosing blocks.

 The state of the calling procedure must be saved so it can resume execution after the call.

 Also saved in a known place is the return address, the location to which the called
routine must transfer after it is finished.

 Finally a jump to the beginning of the code for the called procedure must be generated.

For example, consider the following syntax-directed translation

(1) S  call id ( Elist )

{ for each item p on queue do
emit (‘ param’ p );
emit (‘call’ id.place) }
(2) Elist  Elist , E

{ append E.place to the end of queue }

(3) Elist  E
{ initialize queue to contain only E.place }

 Here, the code for S is the code for Elist, which evaluates the arguments, followed by a
param p statement for each argument, followed by a call statement.

 queue is emptied and then gets a single pointer to the symbol table location for the name
that denotes the value of E.

Simple Type Checker Specification
No ratings yet
Simple Type Checker Specification
24 pages
Module-4 Syntax
No ratings yet
Module-4 Syntax
24 pages
CD Module 4new Full
No ratings yet
CD Module 4new Full
46 pages
Semantic Analysis in Compilers
No ratings yet
Semantic Analysis in Compilers
10 pages
Module 4
No ratings yet
Module 4
24 pages
Compiler Semantic Analysis Guide
No ratings yet
Compiler Semantic Analysis Guide
16 pages
Chapter 4 Semantic Analysis
No ratings yet
Chapter 4 Semantic Analysis
36 pages
Semantic Analysis in Compiler Design
No ratings yet
Semantic Analysis in Compiler Design
35 pages
Semantic Analysis in Compilers
No ratings yet
Semantic Analysis in Compilers
39 pages
UNIT 3 - Chapter 1 in Compiler Design
No ratings yet
UNIT 3 - Chapter 1 in Compiler Design
28 pages
Compiler Design Semantic Analysis
0% (1)
Compiler Design Semantic Analysis
4 pages
Compi Desi CHP 04
No ratings yet
Compi Desi CHP 04
28 pages
Chapter 4 - 6
No ratings yet
Chapter 4 - 6
78 pages
4 Semantic Analysis
No ratings yet
4 Semantic Analysis
20 pages
Compiler Design: Syntax Translation
No ratings yet
Compiler Design: Syntax Translation
247 pages
Chapter 5 Semantic Analysis-SDT
No ratings yet
Chapter 5 Semantic Analysis-SDT
9 pages
Syntax Directed Translation in Compilers
No ratings yet
Syntax Directed Translation in Compilers
43 pages
Module-5-Syntax Directed Translation
No ratings yet
Module-5-Syntax Directed Translation
54 pages
CD Unit 3
No ratings yet
CD Unit 3
23 pages
Chap - 4 - Syntax - Directed - Translation - N07 - G10
No ratings yet
Chap - 4 - Syntax - Directed - Translation - N07 - G10
39 pages
CD Unit-3
No ratings yet
CD Unit-3
36 pages
SE Compiler Chapter 4-SDT
No ratings yet
SE Compiler Chapter 4-SDT
7 pages
Unit 4
No ratings yet
Unit 4
16 pages
Chapter 4
No ratings yet
Chapter 4
43 pages
Module 3
No ratings yet
Module 3
135 pages
Annotated Parse Trees in Compilers
No ratings yet
Annotated Parse Trees in Compilers
60 pages
Chapter 4 Syntax Directed Translation (SDT)
100% (1)
Chapter 4 Syntax Directed Translation (SDT)
6 pages
Chapter 4
No ratings yet
Chapter 4
34 pages
CD Unit 3
No ratings yet
CD Unit 3
18 pages
Unit - IV Syntax-Directed Translation
No ratings yet
Unit - IV Syntax-Directed Translation
50 pages
20-Module 3 - Syntax Directed Definition-21!08!2024
No ratings yet
20-Module 3 - Syntax Directed Definition-21!08!2024
38 pages
Evaluation Order in Syntax-Directed Definitions
No ratings yet
Evaluation Order in Syntax-Directed Definitions
39 pages
Compiler Semantic Analysis Guide
No ratings yet
Compiler Semantic Analysis Guide
46 pages
Applications of Syntax-Directed Translation
No ratings yet
Applications of Syntax-Directed Translation
49 pages
Compiler Design Lec-Four Syntax Directed Translation
No ratings yet
Compiler Design Lec-Four Syntax Directed Translation
29 pages
Unit - 4-Sdd and Sdts - MMM
No ratings yet
Unit - 4-Sdd and Sdts - MMM
47 pages
Syntax-Directed Translation Guide
No ratings yet
Syntax-Directed Translation Guide
51 pages
Ognsemantic Analysis
No ratings yet
Ognsemantic Analysis
68 pages
Lecture 07
No ratings yet
Lecture 07
72 pages
Unit 4 CD
No ratings yet
Unit 4 CD
26 pages
Mod 1 - Syntax Directed Translation
No ratings yet
Mod 1 - Syntax Directed Translation
80 pages
Parsing
No ratings yet
Parsing
6 pages
2 Syntax Directed Transiation
No ratings yet
2 Syntax Directed Transiation
9 pages
Annotated Parse Trees in Compiler Design
No ratings yet
Annotated Parse Trees in Compiler Design
20 pages
CD Unit-4 LM
No ratings yet
CD Unit-4 LM
30 pages
Compiler Design 4
No ratings yet
Compiler Design 4
48 pages
Syntax Directed Translationit
No ratings yet
Syntax Directed Translationit
47 pages
FALLSEM2023-24 BCSE307L TH VL2023240100900 2023-06-06 Reference-Material-I
No ratings yet
FALLSEM2023-24 BCSE307L TH VL2023240100900 2023-06-06 Reference-Material-I
49 pages
Semantic Analysis
No ratings yet
Semantic Analysis
82 pages
CS536 Lec 06
No ratings yet
CS536 Lec 06
25 pages
CC Lec 17
No ratings yet
CC Lec 17
18 pages
Syntax-Directed Translation Guide
No ratings yet
Syntax-Directed Translation Guide
30 pages
Unit 4.1
No ratings yet
Unit 4.1
47 pages
Syntax Directed Translation
No ratings yet
Syntax Directed Translation
47 pages
Syntax-Directed Translations Overview
No ratings yet
Syntax-Directed Translations Overview
52 pages
Compiler Design Chapter-4
100% (2)
Compiler Design Chapter-4
77 pages
Compiler Syntax Translation Guide
No ratings yet
Compiler Syntax Translation Guide
15 pages
Syntax-Directed Translation
No ratings yet
Syntax-Directed Translation
23 pages
EDI Handout
No ratings yet
EDI Handout
3 pages
ASSERTIONS
No ratings yet
ASSERTIONS
14 pages
B.tech Computer Science & Engg Syllabus
No ratings yet
B.tech Computer Science & Engg Syllabus
94 pages
MAT - Series Completion (Number Series, Alphabet Series and Letter Repeating Series)
No ratings yet
MAT - Series Completion (Number Series, Alphabet Series and Letter Repeating Series)
38 pages
AJAX & PHP Question Bank
No ratings yet
AJAX & PHP Question Bank
18 pages
Computer Organization and Architecture LabExperiments
No ratings yet
Computer Organization and Architecture LabExperiments
31 pages
Java Student Record Management System
No ratings yet
Java Student Record Management System
12 pages
Java Programming Overview
100% (1)
Java Programming Overview
76 pages
Final Assignemnt - Part 01 PDF
No ratings yet
Final Assignemnt - Part 01 PDF
2 pages
Suyash Os Notes
No ratings yet
Suyash Os Notes
13 pages
Theory of Computation Exam 2019
No ratings yet
Theory of Computation Exam 2019
3 pages
CE119.O11.2 - LAB02 -Trần Lê Thanh Tùng - 22521621
No ratings yet
CE119.O11.2 - LAB02 -Trần Lê Thanh Tùng - 22521621
7 pages
Paging
No ratings yet
Paging
5 pages
DeepSkin A Deep Learning Approach For Skin Cancer Classification
No ratings yet
DeepSkin A Deep Learning Approach For Skin Cancer Classification
56 pages
5 Fuzzy Expert System
No ratings yet
5 Fuzzy Expert System
6 pages
Reviewer in MATH 5 2nd Periodical Test
100% (4)
Reviewer in MATH 5 2nd Periodical Test
4 pages
AI Algorithms for Pathfinding
No ratings yet
AI Algorithms for Pathfinding
34 pages
Email Spam Classification
No ratings yet
Email Spam Classification
4 pages
C++ Complete Concepts
No ratings yet
C++ Complete Concepts
117 pages
Primary 5 Math Challenge 2023
No ratings yet
Primary 5 Math Challenge 2023
4 pages
BTech R23 II Year Cse Aiml Min
No ratings yet
BTech R23 II Year Cse Aiml Min
40 pages
Automatic Arabic License Plate Recognition
No ratings yet
Automatic Arabic License Plate Recognition
7 pages
Basic Parallel Programming Methods
No ratings yet
Basic Parallel Programming Methods
8 pages
Java Programming Basics Guide
No ratings yet
Java Programming Basics Guide
81 pages
Searching and Sorting Algorithms Overview
No ratings yet
Searching and Sorting Algorithms Overview
23 pages
Mod 5
No ratings yet
Mod 5
21 pages
Analysis and Digital Algorithm Lab Manual
No ratings yet
Analysis and Digital Algorithm Lab Manual
106 pages
Disk Scheduling Algorithms Guide
No ratings yet
Disk Scheduling Algorithms Guide
5 pages
MASP: Scalable Graph-Based Planning Towards Multi-Agent Navigation
No ratings yet
MASP: Scalable Graph-Based Planning Towards Multi-Agent Navigation
8 pages
Chapter 8
No ratings yet
Chapter 8
125 pages

CD Unit 3 - Merged

Uploaded by

CD Unit 3 - Merged

Uploaded by

Unit - 3

SYNTAX DIRECTED TRANSLATION & INTERMEDIATE CODE GENERATION

Syntax directed translation:

Following things are done in Semantic Analysis:

1. Disambiguate Overloaded operators : If an operator is overloaded, one would like to specify

Syntax Directed Deﬁnitions

 Attribute grammar is a medium to provide semantics to the context-free grammar and it

Ex: E → E + T { E.value = E.value + T.value }

The SDD for the above grammar can be written as follow

ANNOTATED PARSE TREE

A --> BCD { C.in = A.in, C.type = B.type }

Computation of Inherited Attributes

The SDD for the above grammar can be written as follow

PRODUCTIONS SEMANTIC RULES

Implementing Syntax Directed Deﬁnitions

Dependency graph for declaration statements.

 S stands for synthesized

Solution: Top down approach:

 First we need to draw the parse tree.

1. Evaluate the expression 4-6/2

Procedure visit (n:node)

for each child m of n, from left to right do

Syntax Directed Translation Schemes (SDT’S)

 It is a CFG with program fragments embedded with the production body.

SDD for infix to postfix translation

GRAMMAR SEMANTIC ACTIONS

Consider again the syntax-directed definition of the desk calculator.

 Type is a notion or a rule that varies from language to language.

Static checking(Semantic check)

 Its helps in detecting and reporting program errors.

 Type of a language construct will be denoted by type expression

Type checking for expressions

Type checking for statements

S id= E { S.type=(if id.type=E.type then void) else type error}

S if E then S1 { S.type=(if E.type=Boolean then S1.type) else type error}

SS1;S2 { S.type=(if S1.type=void and S2.type=void then void) else

Type checking for functions

Benefits of using a machine-independent intermediate form are:

2. A machine-independent code optimizer can be applied to the intermediate representation.

Position of intermediate code generator

parser static intermediate intermediate code

Three ways of intermediate representation:

 Three address code

b uminus b uminus b uminus

(a) Syntax tree (b) Dag

Postfix notation is a linearized representation of a syntax tree; it is a list of the nodes of

a b c uminus * b c uminus * + assign

PRODUCTION SEMANTIC RULE

S  id : = E S.nptr : = mknode(‘assign’,mkleaf(id, id.place), E.nptr)

E  E1 + E2 E.nptr : = mknode(‘+’, E1.nptr, E2.nptr )

E  E1 * E2 E.nptr : = mknode(‘*’, E 1.nptr, E2.nptr )

E  - E1 E.nptr : = mknode(‘uminus’, E1.nptr)

E  id E.nptr : = mkleaf( id, id.place )

Syntax-directed definition to produce syntax trees for assignment statements

Two representations of the syntax tree

Three-address code is a sequence of statements of the general form

where t1 and t2 are compiler-generated temporary names.

 The unraveling of complicated arithmetic expressions and of nested flow-of-control

Three-address code is a linearized representation of a syntax tree or a dag in which

Types of Three-Address Statements:

The common three-address statements are:

1. Assignment statements of the form x : = y op z, where op is a binary arithmetic or logical

2. Assignment instructions of the form x : = op y, where op is a unary operation. Essential unary

3. Copy statements of the form x : = y where the value of y is assigned to x.

7. Indexed assignments of the form x : = y[i] and x[i] : = y.

8. Address and pointer assignments of the form x : = &y , x : = *y, and *x : = y.

Syntax-Directed Translation into Three-Address Code:

Given input a : = b * - c + b * - c, the three-address code is as shown above. The

Syntax-directed definition to produce three-address code for assignments

PRODUCTION SEMANTIC RULES

S  id : = E S.code : = E.code || gen(id.place ‘:=’ E.place)

if E.place = 0 goto S.after

PRODUCTION SEMANTIC RULES

S  while E do S1 S.begin := newlabel;

 The function newtemp returns a sequence of distinct names t 1,t2,….. in response to

Implementation of Three-Address Statements:

8. Address and pointer assignments of the form x : = &y , x : = y, and x : = y.