Major refactoring of the gradle build system.

This commit is contained in:
ghidravore
2019-04-09 11:59:17 -04:00
parent 62a180e0ae
commit f1e50fb079
198 changed files with 2005 additions and 2252 deletions

Binary file not shown.

After

Width:  |  Height:  |  Size: 11 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 4.1 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 7.0 KiB

View File

@@ -0,0 +1,51 @@
/* ###
* IP: GHIDRA
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
/*
WARNING!
This file is copied to all help directories. If you change this file, you must copy it
to each src/main/help/help/shared directory.
Java Help Note: JavaHelp does not accept sizes (like in 'margin-top') in anything but
px (pixel) or with no type marking.
*/
body { margin-bottom: 50px; margin-left: 10px; margin-right: 10px;} /* some padding to improve readability */
li { font-family:times new roman; font-size:14pt; }
h1 { color:#000080; font-family:times new roman; font-size:36pt; font-style:italic; font-weight:bold; text-align:center; }
h2 { margin: 10px; margin-top: 20px; color:#984c4c; font-family:times new roman; font-size:18pt; font-weight:bold; }
h3 { margin-left: 10px; margin-top: 20px; color:#0000ff; font-family:times new roman; font-size:14pt; font-weight:bold; }
h4 { margin-left: 10px; font-family:times new roman; font-size:14pt; font-style:italic; }
/*
P tag code. Most of the help files nest P tags inside of blockquote tags (the was the
way it had been done in the beginning). The net effect is that the text is indented. In
modern HTML we would use CSS to do this. We need to support the Ghidra P tags, nested in
blockquote tags, as well as naked P tags. The following two lines accomplish this. Note
that the 'blockquote p' definition will inherit from the first 'p' definition.
*/
p { margin-left: 40px; font-family:times new roman; font-size:14pt; }
blockquote p { margin-left: 10px; }
p.providedbyplugin { color:#7f7f7f; margin-left: 10px; font-size:14pt; margin-top:100px }
p.ProvidedByPlugin { color:#7f7f7f; margin-left: 10px; font-size:14pt; margin-top:100px }
p.relatedtopic { color:#800080; margin-left: 10px; font-size:14pt; }
p.RelatedTopic { color:#800080; margin-left: 10px; font-size:14pt; }
td { font-family:times new roman; font-size:14pt; vertical-align: top; }
th { font-family:times new roman; font-size:14pt; font-weight:bold; background-color: #EDF3FE; }
code { color: black; font-family: courier new; font-size: 14pt; }

View File

@@ -0,0 +1,323 @@
<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=ISO-8859-1">
<title>Additional P-CODE Operations</title>
<link rel="stylesheet" type="text/css" href="Frontpage.css">
<meta name="generator" content="DocBook XSL Stylesheets V1.78.1">
<link rel="home" href="pcoderef.html" title="P-Code Reference Manual">
<link rel="up" href="pcoderef.html" title="P-Code Reference Manual">
<link rel="prev" href="pseudo-ops.html" title="Pseudo P-CODE Operations">
<link rel="next" href="reference.html" title="Syntax Reference">
</head>
<body bgcolor="white" text="black" link="#0000FF" vlink="#840084" alink="#0000FF">
<div class="navheader">
<table width="100%" summary="Navigation header">
<tr><th colspan="3" align="center">Additional P-CODE Operations</th></tr>
<tr>
<td width="20%" align="left">
<a accesskey="p" href="pseudo-ops.html">Prev</a> </td>
<th width="60%" align="center"> </th>
<td width="20%" align="right"> <a accesskey="n" href="reference.html">Next</a>
</td>
</tr>
</table>
<hr>
</div>
<div class="sect1">
<div class="titlepage"><div><div><h2 class="title" style="clear: both">
<a name="additionalpcode"></a>Additional P-CODE Operations</h2></div></div></div>
<p>
The following opcodes are not generated as part of the raw translation
of a machine instruction into p-code operations, so none of them can be used
in a processor specification. But, they may be
introduced at a later stage by various analysis algorithms.
</p>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="cpui_multiequal"></a>MULTIEQUAL</h3></div></div></div>
<div class="informalexample"><div class="table">
<a name="multiequal.htmltable"></a><table frame="above" width="80%" rules="groups">
<col width="23%">
<col width="15%">
<col width="61%">
<thead><tr>
<td align="center" colspan="2"><span class="bold"><strong>Parameters</strong></span></td>
<td><span class="bold"><strong>Description</strong></span></td>
</tr></thead>
<tbody>
<tr>
<td align="right">input0</td>
<td></td>
<td>Varnode to merge from first basic block.</td>
</tr>
<tr>
<td align="right">input1</td>
<td></td>
<td>Varnode to merge from second basic block.</td>
</tr>
<tr>
<td align="right">[...]</td>
<td></td>
<td>Varnodes to merge from additional basic blocks.</td>
</tr>
<tr>
<td align="right">output</td>
<td></td>
<td>Merged varnode for basic block containing op.</td>
</tr>
</tbody>
<tfoot>
<tr>
<td align="center" colspan="2"><span class="bold"><strong>Semantic statement</strong></span></td>
<td></td>
</tr>
<tr>
<td></td>
<td colspan="2"><span class="emphasis"><em>Cannot be explicitly coded.</em></span></td>
</tr>
</tfoot>
</table>
</div></div>
<p>
This operation represents a copy from one or more possible
locations. From the compiler theory concept of Static Single Assignment form, this is
a <span class="bold"><strong>phi-node</strong></span>. Each input corresponds to a control-flow path
flowing into the basic block containing the <span class="bold"><strong>MULTIEQUAL</strong></span>.
The operator copies a particular input into the output varnode depending on what path
was last executed. All inputs and outputs must be the same size.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="cpui_indirect"></a>INDIRECT</h3></div></div></div>
<div class="informalexample"><div class="table">
<a name="indirect.htmltable"></a><table frame="above" width="80%" rules="groups">
<col width="23%">
<col width="15%">
<col width="61%">
<thead><tr>
<td align="center" colspan="2"><span class="bold"><strong>Parameters</strong></span></td>
<td><span class="bold"><strong>Description</strong></span></td>
</tr></thead>
<tbody>
<tr>
<td align="right">input0</td>
<td></td>
<td>Varnode on which output may depend.</td>
</tr>
<tr>
<td align="right">input1</td>
<td>(<span class="bold"><strong>special</strong></span>)</td>
<td>Code iop of instruction causing effect.</td>
</tr>
<tr>
<td align="right">output</td>
<td></td>
<td>Varnode containing result of effect.</td>
</tr>
</tbody>
<tfoot>
<tr>
<td align="center" colspan="2"><span class="bold"><strong>Semantic statement</strong></span></td>
<td></td>
</tr>
<tr>
<td></td>
<td colspan="2"><span class="emphasis"><em>Cannot be explicitly coded.</em></span></td>
</tr>
</tfoot>
</table>
</div></div>
<p>
An <span class="bold"><strong>INDIRECT</strong></span> operator copies input0 into output,
but the value may be altered in an indirect way
by the operation referred to by input1. The varnode input1 is not part of the
machine state but is really an internal reference to a specific p-code operator that
may be affecting the value of the output varnode. A special address space indicates
input1's use as an internal reference encoding.
An <span class="bold"><strong>INDIRECT</strong></span> op is a placeholder for possible
indirect effects (such as pointer aliasing or missing code) when data-flow
algorithms do not have enough information to follow the data-flow directly. Like
the <span class="bold"><strong>MULTIEQUAL</strong></span>, this op is used
for generating Static Single Assignment form.
</p>
<p>
A constant varnode (zero) for input0 is used by analysis to indicate that the output
of the <span class="bold"><strong>INDIRECT</strong></span> is produced solely by the p-code operation
producing the indirect effect, and there is no possibility that the value existing prior
to the operation was used or preserved.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="cpui_ptradd"></a>PTRADD</h3></div></div></div>
<div class="informalexample"><div class="table">
<a name="ptradd.htmltable"></a><table frame="above" width="80%" rules="groups">
<col width="23%">
<col width="15%">
<col width="61%">
<thead><tr>
<td align="center" colspan="2"><span class="bold"><strong>Parameters</strong></span></td>
<td><span class="bold"><strong>Description</strong></span></td>
</tr></thead>
<tbody>
<tr>
<td align="right">input0</td>
<td></td>
<td>Varnode containing pointer to an array.</td>
</tr>
<tr>
<td align="right">input1</td>
<td></td>
<td>Varnode containing integer index.</td>
</tr>
<tr>
<td align="right">input2</td>
<td>(<span class="bold"><strong>constant</strong></span>)</td>
<td>Constant varnode indicating element size.</td>
</tr>
<tr>
<td align="right">output</td>
<td></td>
<td>Varnode result containing pointer to indexed array entry.</td>
</tr>
</tbody>
<tfoot>
<tr>
<td align="center" colspan="2"><span class="bold"><strong>Semantic statement</strong></span></td>
<td></td>
</tr>
<tr>
<td></td>
<td colspan="2"><span class="emphasis"><em>Cannot be explicitly coded.</em></span></td>
</tr>
</tfoot>
</table>
</div></div>
<p>
This operator serves as a more compact representation of the pointer calculation,
input0 + input1 * input2, but also indicates explicitly that
input0 is a reference to
an array data-type. Input0 is a pointer to the beginning of the array, input1 is an
index into the array, and input2 is a constant indicating the size of
an element in the array. As an
operation, <span class="bold"><strong>PTRADD</strong></span> produces the
pointer value of the element at the indicated index in the array
and stores it in output.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="cpui_ptrsub"></a>PTRSUB</h3></div></div></div>
<div class="informalexample"><div class="table">
<a name="ptrsub.htmltable"></a><table frame="above" width="80%" rules="groups">
<col width="23%">
<col width="15%">
<col width="61%">
<thead><tr>
<td align="center" colspan="2"><span class="bold"><strong>Parameters</strong></span></td>
<td><span class="bold"><strong>Description</strong></span></td>
</tr></thead>
<tbody>
<tr>
<td align="right">input0</td>
<td></td>
<td>Varnode containing pointer to structure.</td>
</tr>
<tr>
<td align="right">input1</td>
<td></td>
<td>Varnode containing integer offset to a subcomponent.</td>
</tr>
<tr>
<td align="right">output</td>
<td></td>
<td>Varnode result containing pointer to the subcomponent.</td>
</tr>
</tbody>
<tfoot>
<tr>
<td align="center" colspan="2"><span class="bold"><strong>Semantic statement</strong></span></td>
<td></td>
</tr>
<tr>
<td></td>
<td colspan="2"><span class="emphasis"><em>Cannot be explicitly coded.</em></span></td>
</tr>
</tfoot>
</table>
</div></div>
<p>
A <span class="bold"><strong>PTRSUB</strong></span> performs the simple pointer calculation,
input0 + input1, but also indicates explicitly that input0 is a
reference to a structured data-type
and one of its subcomponents is being accessed. Input0 is a pointer
to the beginning of the structure, and input1 is a byte offset to the subcomponent.
As an operation, <span class="bold"><strong>PTRSUB</strong></span> produces a
pointer to the subcomponent and stores it in output.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="cpui_cast"></a>CAST</h3></div></div></div>
<div class="informalexample"><div class="table">
<a name="cast.htmltable"></a><table frame="above" width="80%" rules="groups">
<col width="23%">
<col width="15%">
<col width="61%">
<thead><tr>
<td align="center" colspan="2"><span class="bold"><strong>Parameters</strong></span></td>
<td><span class="bold"><strong>Description</strong></span></td>
</tr></thead>
<tbody>
<tr>
<td align="right">input0</td>
<td></td>
<td>Varnode containing value to be copied.</td>
</tr>
<tr>
<td align="right">output</td>
<td></td>
<td>Varnode result of copy.</td>
</tr>
</tbody>
<tfoot>
<tr>
<td align="center" colspan="2"><span class="bold"><strong>Semantic statement</strong></span></td>
<td></td>
</tr>
<tr>
<td></td>
<td colspan="2"><span class="emphasis"><em>Cannot be explicitly coded.</em></span></td>
</tr>
</tfoot>
</table>
</div></div>
<p>
A <span class="bold"><strong>CAST</strong></span> performs identically to the <span class="bold"><strong>COPY</strong></span>
operator but also indicates that there is a forced change in the data-types associated with the varnodes
at this point in the code. The value input0 is strictly copied into output; it is not a conversion cast.
This operator is intended specifically for when the value doesn't change but its
interpretation as a data-type changes at this point.
</p>
</div>
</div>
<div class="navfooter">
<hr>
<table width="100%" summary="Navigation footer">
<tr>
<td width="40%" align="left">
<a accesskey="p" href="pseudo-ops.html">Prev</a> </td>
<td width="20%" align="center"> </td>
<td width="40%" align="right"> <a accesskey="n" href="reference.html">Next</a>
</td>
</tr>
<tr>
<td width="40%" align="left" valign="top">Pseudo P-CODE Operations </td>
<td width="20%" align="center"><a accesskey="h" href="pcoderef.html">Home</a></td>
<td width="40%" align="right" valign="top"> Syntax Reference</td>
</tr>
</table>
</div>
</body>
</html>

View File

@@ -0,0 +1,26 @@
/* ###
* IP: GHIDRA
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
/*
This file contains non-Ghidra style sheet markup. This file will be loaded in addition to
FrontPage.css.
*/
h5 { margin-left: 10px; }
div.informalexample { margin-left: 50px; }
div.example-contents { margin-left: 50px; }
span.term { font-family:times new roman; font-size:14pt; font-weight:bold; }
span.code { font-weight: bold; font-family: courier new; font-size: 14pt; color:#000000; }

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,363 @@
<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=ISO-8859-1">
<title>P-Code Reference Manual</title>
<link rel="stylesheet" type="text/css" href="Frontpage.css">
<link rel="stylesheet" type="text/css" href="languages.css">
<meta name="generator" content="DocBook XSL Stylesheets V1.78.1">
<link rel="home" href="pcoderef.html" title="P-Code Reference Manual">
<link rel="next" href="pcodedescription.html" title="P-Code Operation Reference">
</head>
<body bgcolor="white" text="black" link="#0000FF" vlink="#840084" alink="#0000FF">
<div class="navheader">
<table width="100%" summary="Navigation header">
<tr><th colspan="3" align="center">P-Code Reference Manual</th></tr>
<tr>
<td width="20%" align="left"> </td>
<th width="60%" align="center"> </th>
<td width="20%" align="right"> <a accesskey="n" href="pcodedescription.html">Next</a>
</td>
</tr>
</table>
<hr>
</div>
<div class="article">
<div class="titlepage">
<div>
<div><h1 class="title">
<a name="idm140369391421344"></a>P-Code Reference Manual</h1></div>
<div><p class="releaseinfo">Last updated September 21, 2017</p></div>
</div>
<hr>
</div>
<div class="table">
<a name="mytoc.htmltable"></a><table width="90%" frame="none">
<col width="25%">
<col width="25%">
<col width="25%">
<col width="25%">
<tbody>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_copy" title="COPY">COPY</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_add" title="INT_ADD">INT_ADD</a></td>
<td><a class="link" href="pcodedescription.html#cpui_bool_or" title="BOOL_OR">BOOL_OR</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_load" title="LOAD">LOAD</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_sub" title="INT_SUB">INT_SUB</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float_equal" title="FLOAT_EQUAL">FLOAT_EQUAL</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_store" title="STORE">STORE</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_carry" title="INT_CARRY">INT_CARRY</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float_notequal" title="FLOAT_NOTEQUAL">FLOAT_NOTEQUAL</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_branch" title="BRANCH">BRANCH</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_scarry" title="INT_SCARRY">INT_SCARRY</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float_less" title="FLOAT_LESS">FLOAT_LESS</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_cbranch" title="CBRANCH">CBRANCH</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_sborrow" title="INT_SBORROW">INT_SBORROW</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float_lessequal" title="FLOAT_LESSEQUAL">FLOAT_LESSEQUAL</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_branchind" title="BRANCHIND">BRANCHIND</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_2comp" title="INT_2COMP">INT_2COMP</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float_add" title="FLOAT_ADD">FLOAT_ADD</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_call" title="CALL">CALL</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_negate" title="INT_NEGATE">INT_NEGATE</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float_sub" title="FLOAT_SUB">FLOAT_SUB</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_callind" title="CALLIND">CALLIND</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_xor" title="INT_XOR">INT_XOR</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float_mult" title="FLOAT_MULT">FLOAT_MULT</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pseudo-ops.html#cpui_userdefined" title="USERDEFINED">USERDEFINED</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_and" title="INT_AND">INT_AND</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float_div" title="FLOAT_DIV">FLOAT_DIV</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_return" title="RETURN">RETURN</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_or" title="INT_OR">INT_OR</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float_neg" title="FLOAT_NEG">FLOAT_NEG</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_piece" title="PIECE">PIECE</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_left" title="INT_LEFT">INT_LEFT</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float_abs" title="FLOAT_ABS">FLOAT_ABS</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_subpiece" title="SUBPIECE">SUBPIECE</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_right" title="INT_RIGHT">INT_RIGHT</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float_sqrt" title="FLOAT_SQRT">FLOAT_SQRT</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_int_equal" title="INT_EQUAL">INT_EQUAL</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_sright" title="INT_SRIGHT">INT_SRIGHT</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float_ceil" title="FLOAT_CEIL">FLOAT_CEIL</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_int_notequal" title="INT_NOTEQUAL">INT_NOTEQUAL</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_mult" title="INT_MULT">INT_MULT</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float_floor" title="FLOAT_FLOOR">FLOAT_FLOOR</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_int_less" title="INT_LESS">INT_LESS</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_div" title="INT_DIV">INT_DIV</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float_round" title="FLOAT_ROUND">FLOAT_ROUND</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_int_sless" title="INT_SLESS">INT_SLESS</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_rem" title="INT_REM">INT_REM</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float_nan" title="FLOAT_NAN">FLOAT_NAN</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_int_lessequal" title="INT_LESSEQUAL">INT_LESSEQUAL</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_sdiv" title="INT_SDIV">INT_SDIV</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int2float" title="INT2FLOAT">INT2FLOAT</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_int_slessequal" title="INT_SLESSEQUAL">INT_SLESSEQUAL</a></td>
<td><a class="link" href="pcodedescription.html#cpui_int_srem" title="INT_SREM">INT_SREM</a></td>
<td><a class="link" href="pcodedescription.html#cpui_float2float" title="FLOAT2FLOAT">FLOAT2FLOAT</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_int_zext" title="INT_ZEXT">INT_ZEXT</a></td>
<td><a class="link" href="pcodedescription.html#cpui_bool_negate" title="BOOL_NEGATE">BOOL_NEGATE</a></td>
<td><a class="link" href="pcodedescription.html#cpui_trunc" title="TRUNC">TRUNC</a></td>
</tr>
<tr>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_int_sext" title="INT_SEXT">INT_SEXT</a></td>
<td><a class="link" href="pcodedescription.html#cpui_bool_xor" title="BOOL_XOR">BOOL_XOR</a></td>
<td><a class="link" href="pseudo-ops.html#cpui_cpoolref" title="CPOOLREF">CPOOLREF</a></td>
</tr>
<tr>
<td></td>
<td></td>
<td><a class="link" href="pcodedescription.html#cpui_bool_and" title="BOOL_AND">BOOL_AND</a></td>
<td><a class="link" href="pseudo-ops.html#cpui_new" title="NEW">NEW</a></td>
</tr>
</tbody>
</table>
</div>
<div class="sect1">
<div class="titlepage"><div><div><h2 class="title" style="clear: both">
<a name="index"></a>A Brief Introduction to P-Code</h2></div></div></div>
<p>
P-code is a <span class="emphasis"><em>register transfer language</em></span> designed
for reverse engineering applications. The language is general enough
to model the behavior of many different processors. By modeling in
this way, the analysis of different processors is put into a common
framework, facilitating the development of retargetable analysis
algorithms and applications.
</p>
<p>
Fundamentally, p-code works by translating individual processor instructions
into a sequence of <span class="bold"><strong>p-code operations</strong></span> that take
parts of the processor state as input and output variables
(<span class="bold"><strong>varnodes</strong></span>). The set of unique p-code operations
(distinguished by <span class="bold"><strong>opcode</strong></span>) comprise a fairly tight set
of the arithmetic and logical actions performed by general purpose processors.
The direct translation of instructions into these operations is referred
to as <span class="bold"><strong>raw p-code</strong></span>. Raw p-code can be used to directly emulate
instruction execution and generally follows the same control-flow,
although it may add some of its own internal control-flow. The subset of
opcodes that can occur in raw p-code is described in
<a class="xref" href="pcodedescription.html" title="P-Code Operation Reference">the section called &#8220;P-Code Operation Reference&#8221;</a> and in <a class="xref" href="pseudo-ops.html" title="Pseudo P-CODE Operations">the section called &#8220;Pseudo P-CODE Operations&#8221;</a>, making up
the bulk of this document.
</p>
<p>
P-code is designed specifically to facilitate the
construction of <span class="emphasis"><em>data-flow</em></span> graphs for follow-on analysis of
disassembled instructions. Varnodes and
p-code operators can be thought of explicitly as nodes in these graphs.
Generation of raw p-code is a necessary first step in graph construction,
but additional steps are required, which introduces some new
opcodes. Two of these,
<span class="bold"><strong>MULTIEQUAL</strong></span> and <span class="bold"><strong>INDIRECT</strong></span>,
are specific to the graph construction process, but other opcodes can be introduced during
subsequent analysis and transformation of a graph and help hold recovered data-type relationships.
All of the new opcodes are described in <a class="xref" href="additionalpcode.html" title="Additional P-CODE Operations">the section called &#8220;Additional P-CODE Operations&#8221;</a>, none of which can occur
in the original raw p-code translation. Finally, a few of the p-code operators,
<span class="bold"><strong>CALL</strong></span>,
<span class="bold"><strong>CALLIND</strong></span>, and <span class="bold"><strong>RETURN</strong></span>,
may have their input and output varnodes changed during analysis so that they no
longer match their <span class="emphasis"><em>raw p-code</em></span> form.
</p>
<p>
The core concepts of p-code are:
</p>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140369383722496"></a>Address Space</h3></div></div></div>
<p>
The <span class="bold"><strong>address space</strong></span> for p-code is a generalization
of RAM. It is defined simply as an indexed sequence of bytes that can
be read and written by the p-code operations. For a specific byte, the unique index
that labels it is the byte's <span class="bold"><strong>address</strong></span>. An address space has a
name to identify it, a size that indicates the number of distinct
indices into the space, and an <span class="bold"><strong>endianess</strong></span>
associated with it that indicates how integers and other multi-byte
values are encoded into the space. A typical processor
will have a <span class="bold"><strong>ram</strong></span> space, to model
memory accessible via its main data bus, and
a <span class="bold"><strong>register</strong></span> space for modeling the
processor's general purpose registers. Any data that a processor
manipulates must be in some address space. The specification for a
processor is free to define as many address spaces as it needs. There
is always a special address space, called
a <span class="bold"><strong>constant</strong></span> address space, which is
used to encode any constant values needed for p-code operations. Systems generating
p-code also generally use a dedicated <span class="bold"><strong>temporary</strong></span>
space, which can be viewed as a bottomless source of temporary registers. These
are used to hold intermediate values when modeling instruction behavior.
</p>
<p>
P-code specifications allow the addressable unit of an address
space to be bigger than just a byte. Each address space has
a <span class="bold"><strong>wordsize</strong></span> attribute that can be set
to indicate the number of bytes in a unit. A wordsize which is bigger
than one makes little difference to the representation of p-code. All
the offsets into an address space are still represented internally as
a byte offset. The only exceptions are
the <span class="bold"><strong>LOAD</strong></span> and
<span class="bold"><strong>STORE</strong></span> p-code
operations. These operations read a pointer offset that must be scaled properly to get the
right byte offset when dereferencing the pointer. The wordsize attribute has no effect on
any of the other p-code operations.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140369383712800"></a>Varnode</h3></div></div></div>
<p>
A <span class="bold"><strong>varnode</strong></span> is a generalization of
either a register or a memory location. It is represented by the formal triple:
an address space, an offset into the space, and a size. Intuitively, a
varnode is a contiguous sequence of bytes in some address space that
can be treated as a single value. All manipulation of data by p-code
operations occurs on varnodes.
</p>
<p>
Varnodes by themselves are just a contiguous chunk of bytes,
identified by their address and size, and they have no type. The
p-code operations however can force one of three <span class="emphasis"><em>type</em></span> interpretations
on the varnodes: integer, boolean, and floating-point.
</p>
<div class="informalexample"><div class="itemizedlist"><ul class="itemizedlist compact" style="list-style-type: bullet; ">
<li class="listitem" style="list-style-type: disc">
Operations that manipulate integers always interpret a varnode as a
twos-complement encoding using the endianess associated with the
address space containing the varnode.
</li>
<li class="listitem" style="list-style-type: disc">
A varnode being used as a boolean value is assumed to be a single byte
that can only take the value 0, for <span class="emphasis"><em>false</em></span>, and 1,
for <span class="emphasis"><em>true</em></span>.
</li>
<li class="listitem" style="list-style-type: disc">
Floating-point operations use the encoding expected by the processor being modeled,
which varies depending on the size of the varnode.
For most processors, these encodings are described by the IEEE 754 standard, but
other encodings are possible in principle.
</li>
</ul></div></div>
<p>
</p>
<p>
If a varnode is specified as an offset into
the <span class="bold"><strong>constant</strong></span> address space, that
offset is interpreted as a constant, or immediate value, in any p-code
operation that uses that varnode. The size of the varnode, in this
case, can be treated as the size or precision available for the encoding
of the constant. As with other varnodes, constants only have a type forced
on them by the p-code operations that use them.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140369383704432"></a>P-code Operation</h3></div></div></div>
<p>
A <span class="bold"><strong>p-code operation</strong></span> is the analog of a
machine instruction. All p-code operations have the same basic format
internally. They all take one or more varnodes as input and optionally
produce a single output varnode. The action of the operation is determined by
its <span class="bold"><strong>opcode</strong></span>.
For almost all p-code operations, only the output varnode can have its
value modified; there are no indirect effects of the operation.
The only possible exceptions are <span class="emphasis"><em>pseudo</em></span> operations,
see <a class="xref" href="pseudo-ops.html" title="Pseudo P-CODE Operations">the section called &#8220;Pseudo P-CODE Operations&#8221;</a>, which are sometimes necessary when there
is incomplete knowledge of an instruction's behavior.
</p>
<p>
All p-code operations are associated with the address of the original
processor instruction they were translated from. For a single instruction,
a 1-up counter, starting at zero, is used to enumerate the
multiple p-code operations involved in its translation. The address and
counter as a pair are referred to as the p-code op's
unique <span class="bold"><strong>sequence number</strong></span>. Control-flow of
p-code operations generally follows sequence number order. When execution
of all p-code for one instruction is completed, if the
instruction has <span class="emphasis"><em>fall-through</em></span> semantics, p-code
control-flow picks up with the first p-code operation in sequence corresponding to
the instruction at the fall-through address. Similarly, if a p-code operation
results in a control-flow branch, the first p-code operation in sequence executes
at the destination address.
</p>
<p>
The list of possible
opcodes are similar to many RISC based instruction sets. The effect of
each opcode is described in detail in the following sections,
and a reference table is given
in <a class="xref" href="reference.html" title="Syntax Reference">the section called &#8220;Syntax Reference&#8221;</a>. In general, the size or
precision of a particular p-code operation is determined by the size
of the varnode inputs or output, not by the opcode.
</p>
</div>
</div>
</div>
<div class="navfooter">
<hr>
<table width="100%" summary="Navigation footer">
<tr>
<td width="40%" align="left"> </td>
<td width="20%" align="center"> </td>
<td width="40%" align="right"> <a accesskey="n" href="pcodedescription.html">Next</a>
</td>
</tr>
<tr>
<td width="40%" align="left" valign="top"> </td>
<td width="20%" align="center"> </td>
<td width="40%" align="right" valign="top"> P-Code Operation Reference</td>
</tr>
</table>
</div>
</body>
</html>

View File

@@ -0,0 +1,241 @@
<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=ISO-8859-1">
<title>Pseudo P-CODE Operations</title>
<link rel="stylesheet" type="text/css" href="Frontpage.css">
<link rel="stylesheet" type="text/css" href="languages.css">
<meta name="generator" content="DocBook XSL Stylesheets V1.78.1">
<link rel="home" href="pcoderef.html" title="P-Code Reference Manual">
<link rel="up" href="pcoderef.html" title="P-Code Reference Manual">
<link rel="prev" href="pcodedescription.html" title="P-Code Operation Reference">
<link rel="next" href="additionalpcode.html" title="Additional P-CODE Operations">
</head>
<body bgcolor="white" text="black" link="#0000FF" vlink="#840084" alink="#0000FF">
<div class="navheader">
<table width="100%" summary="Navigation header">
<tr><th colspan="3" align="center">Pseudo P-CODE Operations</th></tr>
<tr>
<td width="20%" align="left">
<a accesskey="p" href="pcodedescription.html">Prev</a> </td>
<th width="60%" align="center"> </th>
<td width="20%" align="right"> <a accesskey="n" href="additionalpcode.html">Next</a>
</td>
</tr>
</table>
<hr>
</div>
<div class="sect1">
<div class="titlepage"><div><div><h2 class="title" style="clear: both">
<a name="pseudo-ops"></a>Pseudo P-CODE Operations</h2></div></div></div>
<p>
Practical analysis systems need to be able to describe operations, whose exact effect on a machine's
memory state is not fully modeled. P-code allows for this by defining a small set of
<span class="emphasis"><em>pseudo</em></span> operators. Such an operator is generally treated as a placeholder
for some, possibly large, sequence of changes to the machine state. In terms of analysis,
either the operator is just carried through as a black-box or it serves as a plug-in point for operator
substitution or other specially tailored transformation. Pseudo operators may violate the requirement
placed on other p-code operations that all effects must be explicit.
</p>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="cpui_userdefined"></a>USERDEFINED</h3></div></div></div>
<div class="informalexample"><div class="table">
<a name="userdefined.htmltable"></a><table frame="above" width="80%" rules="groups">
<col width="23%">
<col width="15%">
<col width="61%">
<thead><tr>
<td align="center" colspan="2"><span class="bold"><strong>Parameters</strong></span></td>
<td><span class="bold"><strong>Description</strong></span></td>
</tr></thead>
<tbody>
<tr>
<td align="right">input0</td>
<td>(<span class="bold"><strong>special</strong></span>)</td>
<td>Constant ID of user-defined op to perform.</td>
</tr>
<tr>
<td align="right">input1</td>
<td></td>
<td>First parameter of user-defined op.</td>
</tr>
<tr>
<td align="right">...</td>
<td></td>
<td>Additional parameters of user-defined op.</td>
</tr>
<tr>
<td align="right">[output]</td>
<td></td>
<td>Optional output of user-defined op.</td>
</tr>
</tbody>
<tfoot>
<tr>
<td align="center" colspan="2"><span class="bold"><strong>Semantic statement</strong></span></td>
<td></td>
</tr>
<tr>
<td></td>
<td colspan="2"><code class="code">userop(input1, ... );</code></td>
</tr>
<tr>
<td></td>
<td colspan="2"><code class="code">output = userop(input1,...);</code></td>
</tr>
</tfoot>
</table>
</div></div>
<p>
This is a placeholder for (a family of) user-definable p-code
instructions. It allows p-code instructions to be defined with
semantic actions that are not fully specified. Machine instructions
that are too complicated or too esoteric to fully implement can use
one or more <span class="bold"><strong>USERDEFINED</strong></span> instructions
as placeholders for their semantics.
</p>
<p>
The first input parameter input0 is a constant ID assigned by the specification
to a particular semantic action. Depending on how the specification
defines the action associated with the ID,
the <span class="bold"><strong>USERDEFINED</strong></span> instruction can take
an arbitrary number of input parameters and optionally have an output
parameter. Exact details are processor and specification dependent.
Ideally, the output parameter is determined by the input
parameters, and no variable is affected except the output
parameter. But this is no longer a strict requirement, side-effects are possible.
Analysis should generally treat these instructions as a &#8220;black-box&#8221; which
still have normal data-flow and can be manipulated symbolically.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="cpui_cpoolref"></a>CPOOLREF</h3></div></div></div>
<div class="informalexample"><div class="table">
<a name="cpoolref.htmltable"></a><table frame="above" width="80%" rules="groups">
<col width="23%">
<col width="15%">
<col width="61%">
<thead><tr>
<td align="center" colspan="2"><span class="bold"><strong>Parameters</strong></span></td>
<td><span class="bold"><strong>Description</strong></span></td>
</tr></thead>
<tbody>
<tr>
<td align="right">input0</td>
<td></td>
<td>Varnode containing pointer offset to object.</td>
</tr>
<tr>
<td align="right">input1</td>
<td>(<span class="bold"><strong>special</strong></span>)</td>
<td>Constant ID indicating type of value to return.</td>
</tr>
<tr>
<td align="right">...</td>
<td></td>
<td>Additional parameters describing value to return.</td>
</tr>
<tr>
<td align="right">output</td>
<td></td>
<td>Varnode to contain requested size, offset, or address.</td>
</tr>
</tbody>
<tfoot>
<tr>
<td align="center" colspan="2"><span class="bold"><strong>Semantic statement</strong></span></td>
<td></td>
</tr>
<tr>
<td></td>
<td colspan="2"><code class="code">output = cpool(input0,intput1);</code></td>
</tr>
</tfoot>
</table>
</div></div>
<p>
This operator returns specific run-time dependent values from the
<span class="emphasis"><em>constant pool</em></span>. This is a concept for object-oriented
instruction sets and other managed code environments, where some details about
how instructions behave can be deferred
until run-time and are not directly encoded in the instruction.
The <span class="bold"><strong>CPOOLREF</strong></span> operator acts a query to the system to
recover this type of information. The first parameter is a
pointer to a specific object, and subsequent parameters are IDs or other special constants
describing exactly what value is requested, relative to the object. The canonical example
is requesting a method address given just an ID describing the method and a specific object, but
<span class="bold"><strong>CPOOLREF</strong></span> can be used as a placeholder for recovering
any important value the system knows about. Details about this instruction, in terms
of emulation and analysis, are necessarily architecture dependent.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="cpui_new"></a>NEW</h3></div></div></div>
<div class="informalexample"><div class="table">
<a name="new.htmltable"></a><table frame="above" width="80%" rules="groups">
<col width="23%">
<col width="15%">
<col width="61%">
<thead><tr>
<td align="center" colspan="2"><span class="bold"><strong>Parameters</strong></span></td>
<td><span class="bold"><strong>Description</strong></span></td>
</tr></thead>
<tbody>
<tr>
<td align="right">input0</td>
<td></td>
<td>Varnode containing class reference</td>
</tr>
<tr>
<td align="right">[input1]</td>
<td></td>
<td>If present, varnode containing count of objects to allocate.</td>
</tr>
<tr>
<td align="right">output</td>
<td></td>
<td>Varnode to contain pointer to allocated memory.</td>
</tr>
</tbody>
<tfoot>
<tr>
<td align="center" colspan="2"><span class="bold"><strong>Semantic statement</strong></span></td>
<td></td>
</tr>
<tr>
<td></td>
<td colspan="2"><code class="code">output = new(input0);</code></td>
</tr>
</tfoot>
</table>
</div></div>
<p>
This operator allocates memory for an object described by the first parameter and
returns a pointer to that memory.
This is used to model object-oriented instruction sets where object allocation is an atomic operation.
Exact details about how memory is affected by a <span class="bold"><strong>NEW</strong></span> operation is generally
not modeled in these cases, so the operator serves as a placeholder to allow analysis to proceed.
</p>
</div>
</div>
<div class="navfooter">
<hr>
<table width="100%" summary="Navigation footer">
<tr>
<td width="40%" align="left">
<a accesskey="p" href="pcodedescription.html">Prev</a> </td>
<td width="20%" align="center"> </td>
<td width="40%" align="right"> <a accesskey="n" href="additionalpcode.html">Next</a>
</td>
</tr>
<tr>
<td width="40%" align="left" valign="top">P-Code Operation Reference </td>
<td width="20%" align="center"><a accesskey="h" href="pcoderef.html">Home</a></td>
<td width="40%" align="right" valign="top"> Additional P-CODE Operations</td>
</tr>
</table>
</div>
</body>
</html>

View File

@@ -0,0 +1,510 @@
<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=ISO-8859-1">
<title>Syntax Reference</title>
<link rel="stylesheet" type="text/css" href="Frontpage.css">
<link rel="stylesheet" type="text/css" href="languages.css">
<meta name="generator" content="DocBook XSL Stylesheets V1.78.1">
<link rel="home" href="pcoderef.html" title="P-Code Reference Manual">
<link rel="up" href="pcoderef.html" title="P-Code Reference Manual">
<link rel="prev" href="additionalpcode.html" title="Additional P-CODE Operations">
</head>
<body bgcolor="white" text="black" link="#0000FF" vlink="#840084" alink="#0000FF">
<div class="navheader">
<table width="100%" summary="Navigation header">
<tr><th colspan="3" align="center">Syntax Reference</th></tr>
<tr>
<td width="20%" align="left">
<a accesskey="p" href="additionalpcode.html">Prev</a> </td>
<th width="60%" align="center"> </th>
<td width="20%" align="right"> </td>
</tr>
</table>
<hr>
</div>
<div class="sect1">
<div class="titlepage"><div><div><h2 class="title" style="clear: both">
<a name="reference"></a>Syntax Reference</h2></div></div></div>
<div class="informalexample"><div class="table">
<a name="ref.htmltable"></a><table width="90%" frame="box" rules="rows">
<col width="25%">
<col width="25%">
<col width="50%">
<thead><tr>
<td><span class="bold"><strong>Name</strong></span></td>
<td><span class="bold"><strong>Syntax</strong></span></td>
<td><span class="bold"><strong>Description</strong></span></td>
</tr></thead>
<tbody>
<tr>
<td>COPY</td>
<td><code class="code">v0 = v1;</code></td>
<td>Copy v1 into v0.</td>
</tr>
<tr>
<td>LOAD</td>
<td>
<div class="table">
<a name="loadtmp.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">* v1</code></td>
</tr>
<tr>
<td><code class="code">*[spc]v1</code></td>
</tr>
<tr>
<td><code class="code">*:2 v1</code></td>
</tr>
<tr>
<td><code class="code">*[spc]:2 v1</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>Dereference v1 as pointer into default space. Optionally specify a space
to load from and size of data in bytes.</td>
</tr>
<tr>
<td>STORE</td>
<td>
<div class="table">
<a name="storetmp.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">*v0 = v1;</code></td>
</tr>
<tr>
<td><code class="code">*[spc]v0 = v1;</code></td>
</tr>
<tr>
<td><code class="code">*:4 v0 = v1;</code></td>
</tr>
<tr>
<td><code class="code">*[spc]:4 v0 = v1;</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>Store in v1 in default space using v0 as pointer. Optionally specify space to store in
and size of data in bytes.</td>
</tr>
<tr>
<td>BRANCH</td>
<td><code class="code">goto v0;</code></td>
<td>Branch execution to address of v0.</td>
</tr>
<tr>
<td>CBRANCH</td>
<td><code class="code">if (v0) goto v1;</code></td>
<td>Branch execution to address of v1 if v0 equals 1 (true).</td>
</tr>
<tr>
<td>BRANCHIND</td>
<td><code class="code">goto [v0];</code></td>
<td>Branch execution to value in v0 viewed as an offset into the current space.</td>
</tr>
<tr>
<td>CALL</td>
<td><code class="code">call v0;</code></td>
<td>Branch execution to address of v0. Hint that the branch is a subroutine call.</td>
</tr>
<tr>
<td>CALLIND</td>
<td><code class="code">call [v0];</code></td>
<td>Branch execution to value in v0 viewed as an offset into the current space.
Hint that the branch is a subroutine call.</td>
</tr>
<tr>
<td>RETURN</td>
<td><code class="code">return [v0];</code></td>
<td>Branch execution to value in v0 viewed as an offset into the current space.
Hint that the branch is a subroutine return.</td>
</tr>
<tr>
<td>INT_EQUAL</td>
<td><code class="code">v0 == v1</code></td>
<td>True if v0 equals v1.</td>
</tr>
<tr>
<td>INT_NOTEQUAL</td>
<td><code class="code">v0 != v1</code></td>
<td>True if v0 does not equal v1.</td>
</tr>
<tr>
<td>INT_SLESS</td>
<td>
<div class="table">
<a name="sless.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">v0 s&lt; v1</code></td>
</tr>
<tr>
<td><code class="code">v1 s&gt; v0</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>True if v0 is less than v1 as a signed integer.</td>
</tr>
<tr>
<td>INT_SLESSEQUAL</td>
<td>
<div class="table">
<a name="slessequal.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">v0 s&lt;= v1</code></td>
</tr>
<tr>
<td><code class="code">v1 s&gt;= v0</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>True if v0 is less than or equal to v1 as a signed integer.</td>
</tr>
<tr>
<td>INT_LESS</td>
<td>
<div class="table">
<a name="less.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">v0 &lt; v1</code></td>
</tr>
<tr>
<td><code class="code">v1 &gt; v0</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>True if v0 is less than v1 as an unsigned integer.</td>
</tr>
<tr>
<td>INT_LESSEQUAL</td>
<td>
<div class="table">
<a name="lessequal.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">v0 &lt;= v1</code></td>
</tr>
<tr>
<td><code class="code">v1 &gt;= v0</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>True if v0 is less than or equal to v1 as an unsigned integer.</td>
</tr>
<tr>
<td>INT_ZEXT</td>
<td><code class="code">zext(v0)</code></td>
<td>Zero extension of v0.</td>
</tr>
<tr>
<td>INT_SEXT</td>
<td><code class="code">sext(v0)</code></td>
<td>Sign extension of v0.</td>
</tr>
<tr>
<td>INT_ADD</td>
<td><code class="code">v0 + v1</code></td>
<td>Addition of v0 and v1 as integers.</td>
</tr>
<tr>
<td>INT_SUB</td>
<td><code class="code">v0 - v1</code></td>
<td>Subtraction of v1 from v0 as integers.</td>
</tr>
<tr>
<td>INT_CARRY</td>
<td><code class="code">carry(v0,v1)</code></td>
<td>True if adding v0 and v1 would produce an unsigned carry.</td>
</tr>
<tr>
<td>INT_SCARRY</td>
<td><code class="code">scarry(v0,v1)</code></td>
<td>True if adding v0 and v1 would produce an signed carry.</td>
</tr>
<tr>
<td>INT_SBORROW</td>
<td><code class="code">sborrow(v0,v1)</code></td>
<td>True if subtracting v1 from v0 would produce a signed borrow.</td>
</tr>
<tr>
<td>INT_2COMP</td>
<td><code class="code">-v0</code></td>
<td>Twos complement of v0.</td>
</tr>
<tr>
<td>INT_NEGATE</td>
<td><code class="code">~v0</code></td>
<td>Bitwise negation of v0.</td>
</tr>
<tr>
<td>INT_XOR</td>
<td><code class="code">v0 ^ v1</code></td>
<td>Bitwise Exclusive Or of v0 with v1.</td>
</tr>
<tr>
<td>INT_AND</td>
<td><code class="code">v0 &amp; v1</code></td>
<td>Bitwise Logical And of v0 with v1.</td>
</tr>
<tr>
<td>INT_OR</td>
<td><code class="code">v0 | v1</code></td>
<td>Bitwise Logical Or of v0 with v1.</td>
</tr>
<tr>
<td>INT_LEFT</td>
<td><code class="code">v0 &lt;&lt; v1</code></td>
<td>Left shift of v0 by v1 bits.</td>
</tr>
<tr>
<td>INT_RIGHT</td>
<td><code class="code">v0 &gt;&gt; v1</code></td>
<td>Unsigned (logical) right shift of v0 by v1 bits.</td>
</tr>
<tr>
<td>INT_SRIGHT</td>
<td><code class="code">v0 s&gt;&gt; v1</code></td>
<td>Signed (arithmetic) right shift of v0 by v1 bits.</td>
</tr>
<tr>
<td>INT_MULT</td>
<td><code class="code">v0 * v1</code></td>
<td>Integer multiplication of v0 and v1.</td>
</tr>
<tr>
<td>INT_DIV</td>
<td><code class="code">v0 / v1</code></td>
<td>Unsigned division of v0 by v1.</td>
</tr>
<tr>
<td>INT_SDIV</td>
<td><code class="code">v0 s/ v1</code></td>
<td>Signed division of v0 by v1.</td>
</tr>
<tr>
<td>INT_REM</td>
<td><code class="code">v0 % v1</code></td>
<td>Unsigned remainder of v0 modulo v1.</td>
</tr>
<tr>
<td>INT_SREM</td>
<td><code class="code">v0 s% v1</code></td>
<td>Signed remainder of v0 modulo v1.</td>
</tr>
<tr>
<td>BOOL_NEGATE</td>
<td><code class="code">!v0</code></td>
<td>Negation of boolean value v0.</td>
</tr>
<tr>
<td>BOOL_XOR</td>
<td><code class="code">v0 ^^ v1</code></td>
<td>Exclusive-Or of booleans v0 and v1.</td>
</tr>
<tr>
<td>BOOL_AND</td>
<td><code class="code">v0 &amp;&amp; v1</code></td>
<td>Logical-And of booleans v0 and v1.</td>
</tr>
<tr>
<td>BOOL_OR</td>
<td><code class="code">v0 || v1</code></td>
<td>Logical-Or of booleans v0 and v1.</td>
</tr>
<tr>
<td>FLOAT_EQUAL</td>
<td><code class="code">v0 f== v1</code></td>
<td>True if v0 equals v1 viewed as floating-point numbers.</td>
</tr>
<tr>
<td>FLOAT_NOTEQUAL</td>
<td><code class="code">v0 f!= v1</code></td>
<td>True if v0 does not equal v1 viewed as floating-point numbers.</td>
</tr>
<tr>
<td>FLOAT_LESS</td>
<td>
<div class="table">
<a name="floatlesstmp.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">v0 f&lt; v1</code></td>
</tr>
<tr>
<td><code class="code">v1 f&gt; v0</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>True if v0 is less than v1 viewed as floating-point numbers.</td>
</tr>
<tr>
<td>FLOAT_LESSEQUAL</td>
<td>
<div class="table">
<a name="floatlessequaltmp.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">v0 f&lt;= v1</code></td>
</tr>
<tr>
<td><code class="code">v1 f&gt;= v0</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>True if v0 is less than or equal to v1 viewed as floating-point numbers.</td>
</tr>
<tr>
<td>FLOAT_NAN</td>
<td><code class="code">nan(v0)</code></td>
<td>True if v0 is not a valid floating-point number (NaN).</td>
</tr>
<tr>
<td>FLOAT_ADD</td>
<td><code class="code">v0 f+ v1</code></td>
<td>Addition of v0 and v1 as floating-point numbers.</td>
</tr>
<tr>
<td>FLOAT_DIV</td>
<td><code class="code">v0 f/ v1</code></td>
<td>Division of v0 by v1 as floating-point numbers.</td>
</tr>
<tr>
<td>FLOAT_MULT</td>
<td><code class="code">v0 f* v1</code></td>
<td>Multiplication of v0 and v1 as floating-point numbers.</td>
</tr>
<tr>
<td>FLOAT_SUB</td>
<td><code class="code">v0 f- v1</code></td>
<td>Subtraction of v1 from v0 as floating-point numbers.</td>
</tr>
<tr>
<td>FLOAT_NEG</td>
<td><code class="code">f- v0</code></td>
<td>Additive inverse of v0 as a floating-point number.</td>
</tr>
<tr>
<td>FLOAT_ABS</td>
<td><code class="code">abs(v0)</code></td>
<td>Absolute value of v0 as a floating-point number.</td>
</tr>
<tr>
<td>FLOAT_SQRT</td>
<td><code class="code">sqrt(v0)</code></td>
<td>Square root of v0 as a floating-point number.</td>
</tr>
<tr>
<td>INT2FLOAT</td>
<td><code class="code">int2float(v0)</code></td>
<td>Floating-point representation of v0 viewed as an integer.</td>
</tr>
<tr>
<td>FLOAT2FLOAT</td>
<td><code class="code">float2float(v0)</code></td>
<td>Copy of floating-point number v0 with more or less precision.</td>
</tr>
<tr>
<td>TRUNC</td>
<td><code class="code">trunc(v0)</code></td>
<td>Signed integer obtained by truncating v0 viewed as a floating-point number.</td>
</tr>
<tr>
<td>FLOAT_CEIL</td>
<td><code class="code">ceil(v0)</code></td>
<td>Nearest integral floating-point value greater than v0, viewed as a floating-point number.</td>
</tr>
<tr>
<td>FLOAT_FLOOR</td>
<td><code class="code">floor(v0)</code></td>
<td>Nearest integral floating-point value less than v0, viewed as a floating-point number.</td>
</tr>
<tr>
<td>FLOAT_ROUND</td>
<td><code class="code">round(v0)</code></td>
<td>Nearest integral floating-point to v0, viewed as a floating-point number.</td>
</tr>
<tr>
<td>SUBPIECE</td>
<td><code class="code">v0:2</code></td>
<td>The least signficant n bytes of v0.</td>
</tr>
<tr>
<td>SUBPIECE</td>
<td><code class="code">v0(2)</code></td>
<td>All but the least significant n bytes of v0.</td>
</tr>
<tr>
<td>PIECE</td>
<td><code class="code">&lt;na&gt;</code></td>
<td>Concatenate two varnodes into a single varnode.</td>
</tr>
<tr>
<td>CPOOLREF</td>
<td><code class="code">cpool(v0,...)</code></td>
<td>Obtain constant pool value.</td>
</tr>
<tr>
<td>NEW</td>
<td>
<div class="table">
<a name="newtmp.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">newobject(v0)</code></td>
</tr>
<tr>
<td><code class="code">newobject(v0,v1)</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>Allocate an object or an array of objects.</td>
</tr>
<tr>
<td>MULTIEQUAL</td>
<td><code class="code">&lt;na&gt;</code></td>
<td>Compiler phi-node: values merging from multiple control-flow paths.</td>
</tr>
<tr>
<td>INDIRECT</td>
<td><code class="code">&lt;na&gt;</code></td>
<td>Indirect effect from input varnode to output varnode.</td>
</tr>
<tr>
<td>CAST</td>
<td><code class="code">&lt;na&gt;</code></td>
<td>Copy from input to output. A hint that the underlying datatype has changed.</td>
</tr>
<tr>
<td>PTRADD</td>
<td><code class="code">&lt;na&gt;</code></td>
<td>Construct a pointer to an element from a pointer to the start of an array and an index.</td>
</tr>
<tr>
<td>PTRSUB</td>
<td><code class="code">&lt;na&gt;</code></td>
<td>Construct a pointer to a field from a pointer to a structure and an offset.</td>
</tr>
</tbody>
</table>
</div></div>
</div>
<div class="navfooter">
<hr>
<table width="100%" summary="Navigation footer">
<tr>
<td width="40%" align="left">
<a accesskey="p" href="additionalpcode.html">Prev</a> </td>
<td width="20%" align="center"> </td>
<td width="40%" align="right"> </td>
</tr>
<tr>
<td width="40%" align="left" valign="top">Additional P-CODE Operations </td>
<td width="20%" align="center"><a accesskey="h" href="pcoderef.html">Home</a></td>
<td width="40%" align="right" valign="top"> </td>
</tr>
</table>
</div>
</body>
</html>

View File

@@ -0,0 +1,438 @@
<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=ISO-8859-1">
<title>SLEIGH</title>
<link rel="stylesheet" type="text/css" href="Frontpage.css">
<link rel="stylesheet" type="text/css" href="languages.css">
<meta name="generator" content="DocBook XSL Stylesheets V1.78.1">
<link rel="home" href="sleigh.html" title="SLEIGH">
<link rel="next" href="sleigh_layout.html" title="2. Basic Specification Layout">
</head>
<body bgcolor="white" text="black" link="#0000FF" vlink="#840084" alink="#0000FF">
<div class="navheader">
<table width="100%" summary="Navigation header">
<tr><th colspan="3" align="center">SLEIGH</th></tr>
<tr>
<td width="20%" align="left"> </td>
<th width="60%" align="center"> </th>
<td width="20%" align="right"> <a accesskey="n" href="sleigh_layout.html">Next</a>
</td>
</tr>
</table>
<hr>
</div>
<div class="article">
<div class="titlepage">
<div>
<div><h1 class="title">
<a name="idm140310883308288"></a>SLEIGH</h1></div>
<div><h3 class="subtitle"><i>A Language for Rapid Processor Specification</i></h3></div>
<div><p class="releaseinfo">Last updated September 1, 2017</p></div>
<div><p class="pubdate">Originally published December 16, 2005</p></div>
</div>
<hr>
</div>
<div class="toc">
<p><b>Table of Contents</b></p>
<dl class="toc">
<dt><span class="sect1"><a href="sleigh.html#idm140310875627168">1. Introduction to P-Code</a></span></dt>
<dd><dl>
<dt><span class="sect2"><a href="sleigh.html#idm140310875617744">1.1. Address Spaces</a></span></dt>
<dt><span class="sect2"><a href="sleigh.html#sleigh_varnodes">1.2. Varnodes</a></span></dt>
<dt><span class="sect2"><a href="sleigh.html#idm140310875600592">1.3. Operations</a></span></dt>
</dl></dd>
<dt><span class="sect1"><a href="sleigh_layout.html">2. Basic Specification Layout</a></span></dt>
<dd><dl>
<dt><span class="sect2"><a href="sleigh_layout.html#idm140310875562464">2.1. Comments</a></span></dt>
<dt><span class="sect2"><a href="sleigh_layout.html#idm140310875560064">2.2. Identifiers</a></span></dt>
<dt><span class="sect2"><a href="sleigh_layout.html#idm140310875558464">2.3. Strings</a></span></dt>
<dt><span class="sect2"><a href="sleigh_layout.html#idm140310875556736">2.4. Integers</a></span></dt>
<dt><span class="sect2"><a href="sleigh_layout.html#idm140310875552544">2.5. White Space</a></span></dt>
</dl></dd>
<dt><span class="sect1"><a href="sleigh_preprocessing.html">3. Preprocessing</a></span></dt>
<dd><dl>
<dt><span class="sect2"><a href="sleigh_preprocessing.html#sleigh_including_files">3.1. Including Files</a></span></dt>
<dt><span class="sect2"><a href="sleigh_preprocessing.html#idm140310875545072">3.2. Preprocessor Macros</a></span></dt>
<dt><span class="sect2"><a href="sleigh_preprocessing.html#idm140310875538656">3.3. Conditional Compilation</a></span></dt>
</dl></dd>
<dt><span class="sect1"><a href="sleigh_definitions.html">4. Basic Definitions</a></span></dt>
<dd><dl>
<dt><span class="sect2"><a href="sleigh_definitions.html#sleigh_endianess_definition">4.1. Endianess Definition</a></span></dt>
<dt><span class="sect2"><a href="sleigh_definitions.html#idm140310875502768">4.2. Alignment Definition</a></span></dt>
<dt><span class="sect2"><a href="sleigh_definitions.html#idm140310875499872">4.3. Space Definitions</a></span></dt>
<dt><span class="sect2"><a href="sleigh_definitions.html#sleigh_naming_registers">4.4. Naming Registers</a></span></dt>
<dt><span class="sect2"><a href="sleigh_definitions.html#idm140310875464736">4.5. Bit Range Registers</a></span></dt>
<dt><span class="sect2"><a href="sleigh_definitions.html#idm140310875451744">4.6. User-Defined Operations</a></span></dt>
</dl></dd>
<dt><span class="sect1"><a href="sleigh_symbols.html">5. Introduction to Symbols</a></span></dt>
<dd><dl>
<dt><span class="sect2"><a href="sleigh_symbols.html#idm140310875423632">5.1. Notes on Namespaces</a></span></dt>
<dt><span class="sect2"><a href="sleigh_symbols.html#sleigh_predefined_symbols">5.2. Predefined Symbols</a></span></dt>
</dl></dd>
<dt><span class="sect1"><a href="sleigh_tokens.html">6. Tokens and Fields</a></span></dt>
<dd><dl>
<dt><span class="sect2"><a href="sleigh_tokens.html#sleigh_defining_tokens">6.1. Defining Tokens and Fields</a></span></dt>
<dt><span class="sect2"><a href="sleigh_tokens.html#idm140310875384864">6.2. Fields as Family Symbols</a></span></dt>
<dt><span class="sect2"><a href="sleigh_tokens.html#idm140310875379232">6.3. Attaching Alternate Meanings to Fields</a></span></dt>
<dt><span class="sect2"><a href="sleigh_tokens.html#sleigh_context_variables">6.4. Context Variables</a></span></dt>
</dl></dd>
<dt><span class="sect1"><a href="sleigh_constructors.html">7. Constructors</a></span></dt>
<dd><dl>
<dt><span class="sect2"><a href="sleigh_constructors.html#idm140310875336416">7.1. The Five Sections of a Constructor</a></span></dt>
<dt><span class="sect2"><a href="sleigh_constructors.html#idm140310875331696">7.2. The Table Header</a></span></dt>
<dt><span class="sect2"><a href="sleigh_constructors.html#sleigh_display_section">7.3. The Display Section</a></span></dt>
<dt><span class="sect2"><a href="sleigh_constructors.html#sleigh_bit_pattern">7.4. The Bit Pattern Section</a></span></dt>
<dt><span class="sect2"><a href="sleigh_constructors.html#sleigh_disassembly_actions">7.5. Disassembly Actions Section</a></span></dt>
<dt><span class="sect2"><a href="sleigh_constructors.html#sleigh_with_block">7.6. The With Block</a></span></dt>
<dt><span class="sect2"><a href="sleigh_constructors.html#sleigh_semantic_section">7.7. The Semantic Section</a></span></dt>
<dt><span class="sect2"><a href="sleigh_constructors.html#sleigh_tables">7.8. Tables</a></span></dt>
<dt><span class="sect2"><a href="sleigh_constructors.html#sleigh_macros">7.9. P-code Macros</a></span></dt>
<dt><span class="sect2"><a href="sleigh_constructors.html#idm140310874869072">7.10. Build Directives</a></span></dt>
<dt><span class="sect2"><a href="sleigh_constructors.html#idm140310874860096">7.11. Delay Slot Directives</a></span></dt>
</dl></dd>
<dt><span class="sect1"><a href="sleigh_context.html">8. Using Context</a></span></dt>
<dd><dl>
<dt><span class="sect2"><a href="sleigh_context.html#idm140310874839872">8.1. Basic Use of Context Variables</a></span></dt>
<dt><span class="sect2"><a href="sleigh_context.html#sleigh_local_change">8.2. Local Context Change</a></span></dt>
<dt><span class="sect2"><a href="sleigh_context.html#sleigh_global_change">8.3. Global Context Change</a></span></dt>
</dl></dd>
<dt><span class="sect1"><a href="sleigh_ref.html">9. P-code Tables</a></span></dt>
</dl>
</div>
<div class="simplesect">
<div class="titlepage"><div><div><h2 class="title" style="clear: both">
<a name="idm140310875635936"></a>History</h2></div></div></div>
<p>
This document describes the syntax for the SLEIGH processor
specification language, which was developed for the GHIDRA
project. The language that is now called SLEIGH has undergone
several redesign iterations, but it can still trace its heritage
from the language SLED, from whom its name is derived. SLED, the
&#8220;Specification Language for Encoding and Decoding&#8221;, was defined by
Norman Ramsey and Mary F. Fernandez as a concise way to define the
translation, in both directions, between machine instructions and
their corresponding assembly statements. This facilitated the
development of architecture independent disassemblers and
assemblers, such as the New Jersey Machine-code Toolkit.
</p>
<p>
The direct predecessor of SLEIGH was an implementation of SLED for
GHIDRA, which concentrated on its reverse-engineering
capabilities. The main addition of SLEIGH is the ability to provide
semantic descriptions of instructions for data-flow and
decompilation analysis. This piece of SLEIGH was originally a
separate language, the Semantic Syntax Language (SSL), very loosely
based on concepts and a language of the same name developed by
Cristina Cifuentes, Mike Van Emmerik and Norman Ramsey, for the
University of Queensland Binary Translator (UQBT) project.
</p>
</div>
<div class="simplesect">
<div class="titlepage"><div><div><h2 class="title" style="clear: both">
<a name="idm140310875632160"></a>Overview</h2></div></div></div>
<p>
SLEIGH is a language for describing the instruction sets of general
purpose microprocessors, in order to facilitate the reverse
engineering of software written for them. SLEIGH was designed for the
GHIDRA reverse engineering platform and is used to describe
microprocessors with enough detail to facilitate two major components
of GHIDRA, the disassembly and decompilation engines. For disassembly,
SLEIGH allows a concise description of the translation from the bit
encoding of machine instructions to human-readable assembly language
statements. Moreover, it does this with enough detail to allow the
disassembly engine to break apart the statement into the mnemonic,
operands, sub-operands, and associated syntax. For decompilation,
SLEIGH describes the translation from machine instructions into
<span class="emphasis"><em>p-code</em></span>. P-code is a Register Transfer Language
(RTL), distinct from SLEIGH, designed to specify
the <span class="emphasis"><em>semantics</em></span> of machine instructions. By
<span class="emphasis"><em>semantics</em></span>, we mean the detailed description of
how an instruction actually manipulates data, in registers and in
RAM. This provides the foundation for the data-flow analysis performed
by the decompiler.
</p>
<p>
A SLEIGH specification typically describes a single microprocessor and
is contained in a single file. The term <span class="emphasis"><em>processor</em></span>
will always refer to this target of the specification.
</p>
<p>
Italics are used when defining terms and for named entities. Bold is used for SLEIGH keywords.
</p>
</div>
<div class="sect1">
<div class="titlepage"><div><div><h2 class="title" style="clear: both">
<a name="idm140310875627168"></a>1. Introduction to P-Code</h2></div></div></div>
<p>
Although p-code is a distinct language from SLEIGH, because a major
purpose of SLEIGH is to specify the translation from machine code to
p-code, this document serves as a primer for p-code. The key concepts
and terminology are presented in this section, and more detail is
given in <a class="xref" href="sleigh_constructors.html#sleigh_semantic_section" title="7.7. The Semantic Section">Section 7.7, &#8220;The Semantic Section&#8221;</a>. There is also a complete set
of tables which list syntax and descriptions for p-code operations in
the Appendix.
</p>
<p>
The design criteria for p-code was to have a language that looks much
like modern assembly instruction sets but capable of modeling any
general purpose processor. Code for different processors can be
translated in a straightforward manner into p-code, and then a single
suite of analysis software can be used to do data-flow analysis and
decompilation. In this way, the analysis software
becomes <span class="emphasis"><em>retargetable</em></span>, and it isn&#8217;t necessary to
redesign it for each new processor being analyzed. It is only
necessary to specify the translation of the processor&#8217;s instruction
set into p-code.
</p>
<p>
So the key properties of p-code are
</p>
<div class="informalexample"><div class="itemizedlist"><ul class="itemizedlist compact" style="list-style-type: bullet; ">
<li class="listitem" style="list-style-type: disc">
The language is machine independent.
</li>
<li class="listitem" style="list-style-type: disc">
The language is designed to model general purpose processors.
</li>
<li class="listitem" style="list-style-type: disc">
Instructions operate on user defined registers and address spaces.
</li>
<li class="listitem" style="list-style-type: disc">
All data is manipulated explicitly. Instructions have no indirect effects.
</li>
<li class="listitem" style="list-style-type: disc">
Individual p-code operations mirror typical processor tasks and concepts.
</li>
</ul></div></div>
<p>
</p>
<p>
SLEIGH is the language which specifies the translation from a machine
instruction to p-code. It specifies both this translation and how to
display the instruction as an assembly statement.
</p>
<p>
A model for a particular processor is built out of three concepts:
the <span class="emphasis"><em>address space</em></span>,
the <span class="emphasis"><em>varnode</em></span>, and
the <span class="emphasis"><em>operation</em></span>. These are generalizations of the
computing concepts of RAM, registers, and machine instructions
respectively.
</p>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875617744"></a>1.1. Address Spaces</h3></div></div></div>
<p>
An <span class="emphasis"><em>address</em></span> space for p-code is a generalization of
the indexed memory (RAM) that a typical processor has access to, and
it is defined simply as an indexed sequence of
memory <span class="emphasis"><em>words</em></span> that can be read and written by
p-code. In almost all cases, a <span class="emphasis"><em>word</em></span> of the space
is a <span class="emphasis"><em>byte</em></span> (8 bits), and we will usually use the
term <span class="emphasis"><em>byte</em></span> instead
of <span class="emphasis"><em>word</em></span>. However, see the discussion of
the <span class="bold"><strong>wordsize</strong></span> attribute of address
spaces below.
</p>
<p>
The defining characteristics of a space are its name and its size. The
size of a space indicates the number of distinct indices into the
space and is usually given as the number of bytes required to encode
an arbitrary index into the space. A space of size 4 requires a 32 bit
integer to specify all indices and contains
2<sup>32</sup> bytes. The index of a byte is usually
referred to as the <span class="emphasis"><em>offset</em></span>, and the offset
together with the name of the space is called
the <span class="emphasis"><em>address</em></span> of the byte.
</p>
<p>
Any manipulation of data that p-code operations perform happens in
some address space. This includes the modeling of data stored in RAM
but also includes the modeling of processor registers. Registers must
be modeled as contiguous sequences of bytes at a specific offset (see
the definition of varnodes below), typically in their own distinct
address space. In order to facilitate the modeling of many different
processors, a SLEIGH specification provides complete control over what
address spaces are defined and where registers are located within
them.
</p>
<p>
Typically, a processor can be modeled with only two spaces,
a <span class="emphasis"><em>ram</em></span> address space that represents the main
memory accessible to the processor via its data-bus, and
a <span class="emphasis"><em>register</em></span> address space that is used to
implement the processor&#8217;s registers. However, the specification
designer can define as many address spaces as needed.
</p>
<p>
There is one address space that is automatically defined for a SLEIGH
specification. This space is used to allocate temporary storage when
the SLEIGH compiler breaks down the expressions describing processor
semantics into individual p-code operations. It is called
the <span class="emphasis"><em>unique</em></span> space. There is also a special address
space, called the <span class="emphasis"><em>const</em></span> space, used as a
placeholder for constant operands of p-code instructions. For the most
part, a SLEIGH specification doesn&#8217;t need to be aware of this space,
but it can be used in certain situations to force values to be
interpreted as constants.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="sleigh_varnodes"></a>1.2. Varnodes</h3></div></div></div>
<p>
A <span class="emphasis"><em>varnode</em></span> is the unit of data manipulated by
p-code. It is simply a contiguous sequence of bytes in some address
space. The two defining characteristics of a varnode are
</p>
<div class="informalexample"><div class="itemizedlist"><ul class="itemizedlist compact" style="list-style-type: bullet; ">
<li class="listitem" style="list-style-type: disc">
The address of the first byte.
</li>
<li class="listitem" style="list-style-type: disc">
The number of bytes (size).
</li>
</ul></div></div>
<p>
With the possible exception of constants treated as varnodes, there is
never any distinction made between one varnode and another. They can
have any size, they can overlap, and any number of them can be
defined.
</p>
<p>
Varnodes by themselves are typeless. An individual p-code operation
forces an interpretation on each varnode that it uses, as either an
integer, a floating-point number, or a boolean value. In the case of
an integer, the varnode is interpreted as having a big endian or
little endian encoding, depending on the specification (see
<a class="xref" href="sleigh_definitions.html#sleigh_endianess_definition" title="4.1. Endianess Definition">Section 4.1, &#8220;Endianess Definition&#8221;</a>). Certain instructions
also distinguish between signed and unsigned interpretations. For a
signed integer, the varnode is considered to have a standard twos
complement encoding. For a boolean interpretation, the varnode must be
a single byte in size. In this special case, the zero encoding of the
byte is considered a <span class="emphasis"><em>false</em></span> value and an encoding
of 1 is a <span class="emphasis"><em>true</em></span> value.
</p>
<p>
These interpretations only apply to the varnode for a particular
operation. A different operation can interpret the same varnode in a
different way. Any consistent meaning assigned to a particular varnode
must be provided and enforced by the specification designer.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875600592"></a>1.3. Operations</h3></div></div></div>
<p>
P-code is intended to emulate a target processor by substituting a
sequence of p-code operations for each machine instruction. Thus every
p-code operation is naturally associated with the address of a
specific machine instruction, but there is usually more than one
p-code operation associated with a single machine instruction. Except
in the case of branching, p-code operations have fall-through control
flow, both within and across machine instructions. For a single
machine instruction, the associated p-code operations execute from
first to last. And if there is no branching, execution picks up with
the first operation corresponding to the next machine instruction.
</p>
<p>
Every p-code operation can take one or more varnodes as input and can
optionally have one varnode as output. The operation can only make a
change to this <span class="emphasis"><em>output varnode</em></span>, which is always indicated
explicitly. Because of this rule, all manipulation of data is
explicit. The operations have no indirect effects. In general, there
is absolutely no restriction on what varnodes can be used as inputs
and outputs to p-code operations. The only exceptions to this are that
constants cannot be used as output varnodes and certain operations
impose restrictions on the <span class="emphasis"><em>size</em></span> of their varnode operands.
</p>
<p>
The actual operations should be familiar to anyone who has studied
general purpose processor instruction sets. They break up into groups.
</p>
<div class="informalexample">
<div class="table">
<a name="ops.htmltable"></a><p class="title"><b>Table 1. P-code Operations</b></p>
<div class="table-contents"><table width="70%" frame="box" rules="all">
<col width="40%">
<col width="60%">
<thead><tr>
<td><span class="bold"><strong>Operation Category</strong></span></td>
<td><span class="bold"><strong>List of Operations</strong></span></td>
</tr></thead>
<tbody>
<tr>
<td>Data Moving</td>
<td><code class="code">COPY, LOAD, STORE</code></td>
</tr>
<tr>
<td>Arithmetic</td>
<td><code class="code">INT_ADD, INT_SUB, INT_CARRY, INT_SCARRY, INT_SBORROW,
INT_2COMP, INT_MULT, INT_DIV, INT_SDIV, INT_REM, INT_SREM</code></td>
</tr>
<tr>
<td>Logical</td>
<td><code class="code">INT_NEGATE, INT_XOR, INT_AND, INT_OR, INT_LEFT, INT_RIGHT, INT_SRIGHT</code></td>
</tr>
<tr>
<td>Integer Comparison</td>
<td><code class="code">INT_EQUAL, INT_NOTEQUAL, INT_SLESS, INT_SLESSEQUAL, INT_LESS, INT_LESSEQUAL</code></td>
</tr>
<tr>
<td>Boolean</td>
<td><code class="code">BOOL_NEGATE, BOOL_XOR, BOOL_AND, BOOL_OR</code></td>
</tr>
<tr>
<td>Floating Point</td>
<td><code class="code">FLOAT_ADD, FLOAT_SUB, FLOAT_MULT, FLOAT_DIV, FLOAT_NEG,
FLOAT_ABS, FLOAT_SQRT, FLOAT_NAN</code></td>
</tr>
<tr>
<td>Floating Point Compare</td>
<td><code class="code">FLOAT_EQUAL, FLOAT_NOTEQUAL, FLOAT_LESS, FLOAT_LESSEQUAL</code></td>
</tr>
<tr>
<td>Floating Point Conversion</td>
<td><code class="code">INT2FLOAT, FLOAT2FLOAT, TRUNC, CEIL, FLOOR, ROUND</code></td>
</tr>
<tr>
<td>Branching</td>
<td><code class="code">BRANCH, CBRANCH, BRANCHIND, CALL, CALLIND, RETURN</code></td>
</tr>
<tr>
<td>Extension/Truncation</td>
<td><code class="code">INT_ZEXT, INT_SEXT, PIECE, SUBPIECE</code></td>
</tr>
<tr>
<td>Managed Code</td>
<td><code class="code">CPOOLREF, NEW</code></td>
</tr>
</tbody>
</table></div>
</div>
<br class="table-break">
</div>
<p>
We postpone a full discussion of the individual operations until <a class="xref" href="sleigh_constructors.html#sleigh_semantic_section" title="7.7. The Semantic Section">Section 7.7, &#8220;The Semantic Section&#8221;</a>.
</p>
</div>
</div>
</div>
<div class="navfooter">
<hr>
<table width="100%" summary="Navigation footer">
<tr>
<td width="40%" align="left"> </td>
<td width="20%" align="center"> </td>
<td width="40%" align="right"> <a accesskey="n" href="sleigh_layout.html">Next</a>
</td>
</tr>
<tr>
<td width="40%" align="left" valign="top"> </td>
<td width="20%" align="center"> </td>
<td width="40%" align="right" valign="top"> 2. Basic Specification Layout</td>
</tr>
</table>
</div>
</body>
</html>

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,363 @@
<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=ISO-8859-1">
<title>8. Using Context</title>
<link rel="stylesheet" type="text/css" href="Frontpage.css">
<link rel="stylesheet" type="text/css" href="languages.css">
<meta name="generator" content="DocBook XSL Stylesheets V1.78.1">
<link rel="home" href="sleigh.html" title="SLEIGH">
<link rel="up" href="sleigh.html" title="SLEIGH">
<link rel="prev" href="sleigh_constructors.html" title="7. Constructors">
<link rel="next" href="sleigh_ref.html" title="9. P-code Tables">
</head>
<body bgcolor="white" text="black" link="#0000FF" vlink="#840084" alink="#0000FF">
<div class="navheader">
<table width="100%" summary="Navigation header">
<tr><th colspan="3" align="center">8. Using Context</th></tr>
<tr>
<td width="20%" align="left">
<a accesskey="p" href="sleigh_constructors.html">Prev</a> </td>
<th width="60%" align="center"> </th>
<td width="20%" align="right"> <a accesskey="n" href="sleigh_ref.html">Next</a>
</td>
</tr>
</table>
<hr>
</div>
<div class="sect1">
<div class="titlepage"><div><div><h2 class="title" style="clear: both">
<a name="sleigh_context"></a>8. Using Context</h2></div></div></div>
<p>
For most practical specifications, the disassembly and semantic
meaning of an instruction can be determined by looking only at the
bits in the encoding of that instruction. SLEIGH syntax reflects this
fact as every constructor or attached register is ultimately decided
by examining <span class="emphasis"><em>fields</em></span>, the syntactic representation
of these instruction bits. In some cases however, the instruction
encoding itself may not be enough. Additional information, which we
refer to as <span class="emphasis"><em>context</em></span>, may be necessary to fully
resolve the meaning of the instruction.
</p>
<p>
In truth, almost every modern processor has multiple modes of
operation, where the exact interpretation of instructions may depend
on that mode. Typical examples include distinguishing between a 16-bit
mode and a 32-bit mode, or between a big endian mode or a little
endian mode. But for the specification designer, these are of little
consequence because most software for such a processor will run in
only one mode without ever changing it. The designer need only pick
the most popular or most important mode for his projects and design to
that. If there is in fact software that runs under a different mode,
the other mode can be described in a separate specification.
</p>
<p>
However, for certain processors or software, the need to distinguish
between different interpretations of the same instruction encoding,
based on context, may be a crucial part of the disassembly and
analysis process. There are two typical situations where this becomes
necessary.
</p>
<div class="informalexample"><div class="itemizedlist"><ul class="itemizedlist compact" style="list-style-type: bullet; ">
<li class="listitem" style="list-style-type: disc">
The processor supports two (or more) separate instruction
sets. The set to use is usually determined by special bits in a status
register, and a single piece of software frequently switches between
these modes.
</li>
<li class="listitem" style="list-style-type: disc">
The processor supports instructions that temporarily affect
the execution of the immediately following instruction(s). For
example, many processors support hardware <span class="emphasis"><em>loop</em></span> instructions that
automatically cause the following instructions to repeat without an
explicit instruction causing the branching and loop counting.
</li>
</ul></div></div>
<p>
</p>
<p>
SLEIGH solves these problems by introducing <span class="emphasis"><em>context
variables</em></span>. The syntax for defining these symbols was
described in <a class="xref" href="sleigh_tokens.html#sleigh_context_variables" title="6.4. Context Variables">Section 6.4, &#8220;Context Variables&#8221;</a>. As mentioned
there, the easiest and most common way to use a context variable is as
just another field to use in our bit patterns. It gives us the extra
information we need to distinguish between different instructions
whose encodings are otherwise the same.
</p>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310874839872"></a>8.1. Basic Use of Context Variables</h3></div></div></div>
<p>
Suppose a processor supports the use of two different sets of
registers in its main addressing mode, based on the setting of a
status bit which can be changed dynamically. If an instruction is
executed with this bit cleared, then one set of registers is used, and
if the bit is set, the other registers are used. The instructions
otherwise behave identically.
</p>
<div class="informalexample"><pre class="programlisting">
define endian=big;
define space ram type=ram_space size=4 default;
define space register type=register_space size=4;
define register offset=0 size=4 [ r0 r1 r2 r3 r4 r5 r6 r7 ];
define register offset=0x100 size=4 [ s0 s1 s2 s3 s4 s5 s6 s7 ];
define token instr(16)
op=(10,15) rreg1=(7,9) sreg1=(7,9) imm=(0,6)
;
define context statusreg
mode=(3,3)
;
attach variables [ rreg1 ] [ r0 r1 r2 r3 r4 r5 r6 r7 ];
attach variables [ sreg1 ] [ s0 s1 s2 s3 s4 s5 s6 s7 ];
Reg1: rreg1 is mode=0 &amp; rreg1 { export rreg1; }
Reg1: sreg1 is mode=1 &amp; sreg1 { export sreg1; }
:addi Reg1,#imm is op=1 &amp; Reg1 &amp; imm { Reg1 = Reg1 + imm; }
</pre></div>
<p>
</p>
<p>
In this example the symbol <span class="emphasis"><em>Reg1</em></span> uses the 3 bits
(7,9) to select one of eight registers. If the context
variable <span class="emphasis"><em>mode</em></span> is set to 0, it selects
an <span class="emphasis"><em>r</em></span> register, through
the <span class="emphasis"><em>rreg1</em></span> field. If <span class="emphasis"><em>mode</em></span> is
set to 1 on the other hand, an <span class="emphasis"><em>s</em></span> register is
selected instead
via <span class="emphasis"><em>sreg1</em></span>. The <span class="emphasis"><em>addi</em></span>
instruction (encoded as 0x0590 for example) can disassemble in one of
two ways.
</p>
<div class="informalexample"><pre class="programlisting">
addi r3,#0x10 <span class="bold"><strong>OR</strong></span>
addi s3,#0x10
</pre></div>
<p>
</p>
<p>
This is the same behavior as if <span class="emphasis"><em>mode</em></span> were defined
as a field instead of a context variable, except that there is nothing
in the instruction encoding itself which indicates which of the two
forms will be chosen. An engine doing the disassembly will have global
state associated with the <span class="emphasis"><em>mode</em></span> variable that will
make the final decision about which form to generate. The setting of
this state is (at least partially) out of the control of SLEIGH,
although see the following sections.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="sleigh_local_change"></a>8.2. Local Context Change</h3></div></div></div>
<p>
SLEIGH can make direct modifications to context variables through
statements in the disassembly action section of a constructor. The
left-hand side of an assignment statement in this section can be a context variable,
see <a class="xref" href="sleigh_constructors.html#sleigh_general_actions" title="7.5.2. General Actions and Pattern Expressions">Section 7.5.2, &#8220;General Actions and Pattern Expressions&#8221;</a>. Because the result of this
assignment is calculated in the middle of the instruction disassembly,
the change in value of the context variable can potentially affect any
remaining parsing for that instruction. A modal variable is being
added to what was otherwise a stateless grammar, a common technique in
many practical parsing engines.
</p>
<p>
Any assignment statement changing a context variable is immediately
executed upon the successful match of the constructor containing the
statement and can be used to guide the parsing of the constructor's
operands. We introduce two more instructions to the example
specification from the previous section.
</p>
<div class="informalexample"><pre class="programlisting">
:raddi Reg1,#imm is op=2 &amp; Reg1 &amp; imm [ mode=0; ] {
Reg1 = Reg1 + imm;
}
:saddi Reg1,#imm is op=3 &amp; Reg1 &amp; imm [ mode=1; ] {
Reg1 = Reg1 + imm;
}
</pre></div>
<p>
</p>
<p>
Notice that both new constructors modify the context
variable <span class="emphasis"><em>mode</em></span>. The raddi instruction sets mode to
0 and effectively guarantees that an <span class="emphasis"><em>r</em></span> register
will be produced by the disassembly. Similarly,
the <span class="emphasis"><em>saddi</em></span> instruction can force
an <span class="emphasis"><em>s</em></span> register. Both are in contrast to
the <span class="emphasis"><em>addi</em></span> instruction, which depends on a global
state. The changes to <span class="emphasis"><em>mode</em></span> made by these
instructions only persist for parsing of that single instruction. For
any following instructions, if the matching constructors
use <span class="emphasis"><em>mode</em></span>, its value will have reverted to its
original global state. The same holds for any context variable
modified with this syntax. If an instruction needs to permanently
modify the state of a context variable, the designer must use
constructions described in <a class="xref" href="sleigh_context.html#sleigh_global_change" title="8.3. Global Context Change">Section 8.3, &#8220;Global Context Change&#8221;</a>.
</p>
<p>
Clearly, the behavior of the above example could be easily replicated
without using context variables at all and having the selection of a
register set simply depend directly on the <span class="emphasis"><em>op</em></span>
field. But, with more complicated addressing modes, local modification
of context variables can drastically reduce the complexity and size of
a specification.
</p>
<p>
At the point where a modification is made to a context variable, the
specification designer has the guarantee that none of the operands of
the constructor have been evaluated yet, so if their matching depends
on this context variable, they will be affected by the change. In
contrast, the matching of any ancestor constructor cannot be
affected. Other constructors, which are not direct ancestors or
descendants, may or may not be affected by the change, depending on
the order of evaluation. It is usually best not to depend on this
ordering when designing the specification, with the possible exception
of orderings which are guaranteed
by <span class="bold"><strong>build</strong></span> directives.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="sleigh_global_change"></a>8.3. Global Context Change</h3></div></div></div>
<p>
It is possible for an instruction to attempt a permanent change to a
context variable, which would then affect the parsing of other
instructions, by using the <span class="bold"><strong>globalset</strong></span>
directive in a disassembly action. As mentioned in the previous
section, context variables have an associated global state, which can
be used during constructor matching. A complete model for this state
is, unfortunately, outside the scope of SLEIGH. The disassembly engine
has to make too many decisions about what is getting disassembled and
what assumptions are being made to give complete control of the
context to SLEIGH. Because of this caveat, SLEIGH syntax for making
permanent context changes should be viewed as a suggestion to the
disassembly engine.
</p>
<p>
For processors that support multiple modes, there are typically
specific instructions that switch between these modes. Extending the
example from the previous sections, we add two instructions to the
specification for permanently switching which register set is being
used.
</p>
<div class="informalexample"><pre class="programlisting">
:rmode is op=32 &amp; rreg1=0 &amp; imm=0
[ mode=0; globalset(inst_next,mode); ]
{}
:smode is op=33 &amp; rreg1=0 &amp; imm=0
[ mode=1; globalset(inst_next,mode); ]
{}
</pre></div>
<p>
</p>
<p>
The register set is, as before, controlled by
the <span class="emphasis"><em>mode</em></span> variable, and as with a local change to
context, the variable is assigned to inside the square
brackets. The <span class="emphasis"><em>rmode</em></span> instruction
sets <span class="emphasis"><em>mode</em></span> to 0, in order to
select <span class="emphasis"><em>r</em></span> registers
via <span class="emphasis"><em>rreg1</em></span>, and <span class="emphasis"><em>smode</em></span>
sets <span class="emphasis"><em>mode</em></span> to 1 in order to
select <span class="emphasis"><em>s</em></span> registers. As is described in
<a class="xref" href="sleigh_context.html#sleigh_local_change" title="8.2. Local Context Change">Section 8.2, &#8220;Local Context Change&#8221;</a>, these assignments by themselves
cause only a local context change. However, the
subsequent <span class="bold"><strong>globalset</strong></span> directives make
the change persist outside of the the instructions
themselves. The <span class="bold"><strong>globalset</strong></span> directive
takes two parameters, the second being the particular context variable
being changed. The first parameter indicates the first address where
the new context takes effect. In the example, the expectation is that
a mode change affects any subsequent instructions. So the first
parameter to <span class="bold"><strong>globalset</strong></span> here
is <span class="emphasis"><em>inst_next</em></span>, indicating that the new value
of <span class="emphasis"><em>mode</em></span> begins at the next address.
</p>
<div class="sect3">
<div class="titlepage"><div><div><h4 class="title">
<a name="sleigh_contextflow"></a>8.3.1. Context Flow</h4></div></div></div>
<p>
A global change to context that affects instruction decoding is typically
open-ended. I.e. once the mode switching instruction is executed, a permanent change
is made to the run-time processor state, and all future instruction decoding is
affected, until another mode switch is encountered. In terms of SLEIGH by default,
the effect of a <span class="bold"><strong>globalset</strong></span> directive
follows <span class="emphasis"><em>flow</em></span>. Starting from the address specified in the directive,
the change in context follows the control-flow of the instructions, through
branches and calls, until an execution path terminates or another context change
is encountered.
</p>
<p>
Flow following behavior can be overridden by adding the <span class="bold"><strong>noflow</strong></span>
attribute to the definition of the context field. (See <a class="xref" href="sleigh_tokens.html#sleigh_context_variables" title="6.4. Context Variables">Section 6.4, &#8220;Context Variables&#8221;</a>)
In this case, a <span class="bold"><strong>globalset</strong></span> directive only affects the context
of a single instruction at the specified address. Subsequent instructions
retain their original context. This can be useful in a variety of situations but is typically
used to let one instruction alter the behavior, not necessarily the decoding,
of the following instruction. In the example below,
an indirect branch instruction jumps through a link register <span class="emphasis"><em>lr</em></span>. If the previous
instruction moves the program counter in to <span class="emphasis"><em>lr</em></span>, it communicates this to the
branch instruction through the <span class="emphasis"><em>LRset</em></span> context variable so that the branch can
be interpreted as a return, rather than a generic indirect branch.
</p>
<div class="informalexample"><pre class="programlisting">
define context contextreg
LRset = (1,1) noflow # 1 if the instruction before was a mov lr,pc
;
<span class="weak">...</span>
mov lr,pc is opcode=34 &amp; lr &amp; pc
[ LRset=1; globalset(inst_next,LRset); ] { lr = pc; }
<span class="weak">...</span>
blr is opcode=35 &amp; reg=15 &amp; LRset=0 { goto [lr]; }
blr is opcode=35 &amp; reg=15 &amp; LRset=1 { return [lr]; }
</pre></div>
<p>
</p>
<p>
An alternative to the <span class="bold"><strong>noflow</strong></span> attribute is to simply issue
multiple directives within a single constructor, so an explicit end to a context change
can be given. The value of the variable exported to the global state
is the one in affect at the point where the directive is issued. Thus,
after one <span class="bold"><strong>globalset</strong></span>, the same context
variable can be assigned a different value, followed by
another <span class="bold"><strong>globalset</strong></span> for a different
address.
</p>
<p>
Because context in SLEIGH is controlled by a disassembly process,
there are some basic caveats to the use of
the <span class="bold"><strong>globalset</strong></span> directive. With
<span class="emphasis"><em>flowing</em></span> context changes,
there is no guarantee of what global state will be in affect at a
particular address. During disassembly, at any given
point, the process may not have uncovered all the relevant directives,
and the known directives may not necessarily be consistent. In
general, for most processors, the disassembly at a particular address
is intended to be absolute. So given enough information, it should be
possible to make a definitive determination of what the context is at
a certain address, but there is no guarantee. It is up to the
disassembly process to fully determine where context changes begin and
end and what to do if there are conflicts.
</p>
</div>
</div>
</div>
<div class="navfooter">
<hr>
<table width="100%" summary="Navigation footer">
<tr>
<td width="40%" align="left">
<a accesskey="p" href="sleigh_constructors.html">Prev</a> </td>
<td width="20%" align="center"> </td>
<td width="40%" align="right"> <a accesskey="n" href="sleigh_ref.html">Next</a>
</td>
</tr>
<tr>
<td width="40%" align="left" valign="top">7. Constructors </td>
<td width="20%" align="center"><a accesskey="h" href="sleigh.html">Home</a></td>
<td width="40%" align="right" valign="top"> 9. P-code Tables</td>
</tr>
</table>
</div>
</body>
</html>

View File

@@ -0,0 +1,353 @@
<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=ISO-8859-1">
<title>4. Basic Definitions</title>
<link rel="stylesheet" type="text/css" href="Frontpage.css">
<link rel="stylesheet" type="text/css" href="languages.css">
<meta name="generator" content="DocBook XSL Stylesheets V1.78.1">
<link rel="home" href="sleigh.html" title="SLEIGH">
<link rel="up" href="sleigh.html" title="SLEIGH">
<link rel="prev" href="sleigh_preprocessing.html" title="3. Preprocessing">
<link rel="next" href="sleigh_symbols.html" title="5. Introduction to Symbols">
</head>
<body bgcolor="white" text="black" link="#0000FF" vlink="#840084" alink="#0000FF">
<div class="navheader">
<table width="100%" summary="Navigation header">
<tr><th colspan="3" align="center">4. Basic Definitions</th></tr>
<tr>
<td width="20%" align="left">
<a accesskey="p" href="sleigh_preprocessing.html">Prev</a> </td>
<th width="60%" align="center"> </th>
<td width="20%" align="right"> <a accesskey="n" href="sleigh_symbols.html">Next</a>
</td>
</tr>
</table>
<hr>
</div>
<div class="sect1">
<div class="titlepage"><div><div><h2 class="title" style="clear: both">
<a name="sleigh_definitions"></a>4. Basic Definitions</h2></div></div></div>
<p>
SLEIGH files must start with all the definitions needed by the rest of
the specification. All definition statements start with the keyword
<span class="bold"><strong>define</strong></span> and end with a semicolon &#8216;;&#8217;.
</p>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="sleigh_endianess_definition"></a>4.1. Endianess Definition</h3></div></div></div>
<p>
The first definition in any SLEIGH specification must be for endianess. Either
</p>
<div class="informalexample"><pre class="programlisting">
define endian=big; <span class="emphasis"><em>OR</em></span>
define endian=little;
</pre></div>
<p>
This defines how the processor interprets contiguous sequences of
bytes as integers. It effects how integer fields within an instruction
are interpreted (see <a class="xref" href="sleigh_tokens.html#sleigh_defining_tokens" title="6.1. Defining Tokens and Fields">Section 6.1, &#8220;Defining Tokens and Fields&#8221;</a>), and
it also effects the details of how the processor is supposed to
implement atomic operations like integer addition and integer
compare. The specification designer should only need to worry about
these details when labeling instruction fields, otherwise the
specification language will hide endianess issues.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875502768"></a>4.2. Alignment Definition</h3></div></div></div>
<p>
An alignment definition looks like
</p>
<div class="informalexample"><pre class="programlisting">
define alignment=<span class="bold"><strong>integer</strong></span>;
</pre></div>
<p>
This specifies the byte alignment of instructions within their address
space. It defaults to 1 or no alignment. When disassembling an
instruction at a particular, the disassembler checks the alignment of
the address against this value and can opt to flag an unaligned
instruction as an error.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875499872"></a>4.3. Space Definitions</h3></div></div></div>
<p>
The definition of an address space looks like
</p>
<div class="informalexample"><pre class="programlisting">
define space <span class="bold"><strong>spacename attributes</strong></span> ;
</pre></div>
<p>
The <span class="emphasis"><em>spacename</em></span> is the name of the new space,
and <span class="emphasis"><em>attributes</em></span> looks like zero or more of the
following lines:
</p>
<div class="informalexample"><pre class="programlisting">
type=(ram_space|rom_space|register_space)
size=<span class="bold"><strong>integer</strong></span>
default
wordsize=<span class="bold"><strong>integer</strong></span>
</pre></div>
<p>
The only required attribute is <span class="bold"><strong>size</strong></span>
which specifies the number of bytes needed to address any byte within
the space, for example a 32-bit address space has size 4.
</p>
<p>
A space of type <span class="bold"><strong>ram_space</strong></span> is defined as follows:
</p>
<div class="informalexample"><div class="itemizedlist"><ul class="itemizedlist compact" style="list-style-type: bullet; ">
<li class="listitem" style="list-style-type: disc">
It is read/write.
</li>
<li class="listitem" style="list-style-type: disc">
It is part of the standard memory map of the processor.
</li>
<li class="listitem" style="list-style-type: disc">
It is addressable in the sense that the processor may load
and store from dynamic pointers into the space.
</li>
</ul></div></div>
<p>
</p>
<p>
A space of type <span class="bold"><strong>register_space</strong></span> is
intended to model the processor&#8217;s general-purpose registers. In terms
of accessing and manipulating data within the space, SLEIGH and p-code
make no distinction between the
type <span class="bold"><strong>ram_space</strong></span> or the
type <span class="bold"><strong>register_space</strong></span>. But there are
still some distinguishing properties of a space labeled
with <span class="bold"><strong>register_space</strong></span>.
</p>
<div class="informalexample"><div class="itemizedlist"><ul class="itemizedlist compact" style="list-style-type: bullet; ">
<li class="listitem" style="list-style-type: disc">
It is read/write.
</li>
<li class="listitem" style="list-style-type: disc">
It is <span class="emphasis"><em>not</em></span> part of the standard memory map of the processor.
</li>
<li class="listitem" style="list-style-type: disc">
In terms of GHIDRA, there will not be separate windows for
the space and references into the space will not be stored.
</li>
<li class="listitem" style="list-style-type: disc">
Named symbols within the space will have Register objects
associated with them in GHIDRA.
</li>
<li class="listitem" style="list-style-type: disc">
It is <span class="emphasis"><em>not</em></span> addressable. Data-flow
analysis will assume that data within the space cannot be
manipulated indirectly via pointer, so there is no pointer
aliasing. Make sure this is true!
</li>
</ul></div></div>
<p>
</p>
<p>
A space of type <span class="bold"><strong>rom_space</strong></span> has seen
little use so far but is intended to be the same as
a <span class="bold"><strong>ram_space</strong></span> that is not writable.
</p>
<p>
At least one space needs to be labeled with
the <span class="bold"><strong>default</strong></span> attribute. This should be
the space that the processor accesses with its main address bus. In
terms of the rest of the specification file, this sets the default
space referred to by the &#8216;*&#8217; operator (see
<a class="xref" href="sleigh_constructors.html#sleigh_star_operator" title="7.7.1.2. The '*' Operator">Section 7.7.1.2, &#8220;The '*' Operator&#8221;</a>). It also has meaning to
GHIDRA.
</p>
<p>
The average 32-bit processor requires only the following two space definitions.
</p>
<div class="informalexample"><pre class="programlisting">
define space ram type=ram_space size=4 default;
define space register type=register_space size=4;
</pre></div>
<p>
The <span class="bold"><strong>wordsize</strong></span> attribute can be used to
specify the size of the memory location referred to with a single
address. If a space has <span class="bold"><strong>wordsize</strong></span> two,
then each address of the space refers to 16 bits of data, rather than
8 bits. If the space has <span class="bold"><strong>size</strong></span> two,
then there are still 2<sup>16</sup> different
addresses, but since each address accesses two bytes, there are twice
as many bytes, 2<sup>17</sup>, in the space. If
the <span class="bold"><strong>wordsize</strong></span> attribute is not
specified, the size of a memory location defaults to one byte (8
bits).
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="sleigh_naming_registers"></a>4.4. Naming Registers</h3></div></div></div>
<p>
The general purpose registers of the processors can be named with the
following define syntax:
</p>
<div class="informalexample"><pre class="programlisting">
define <span class="bold"><strong>spacename</strong></span> offset=<span class="bold"><strong>integer</strong></span> size=<span class="bold"><strong>integer stringlist</strong></span> ;
</pre></div>
<p>
A <span class="emphasis"><em>stringlist</em></span> is either a single string or a white
space separated list of strings in square brackets &#8216;[&#8217; and &#8216;]&#8217;. A
string of just &#8220;_&#8221; indicates a skip in the sequence for that
definition. The offset corresponding to that position in the list of
names will not have a varnode defined at it.
</p>
<p>
This defines specific varnodes within the indicated address
space. Each name in the list is assigned to a varnode in turn starting
at the indicated offset within the space. Each varnode occupies the
indicated number of bytes in size. There is no restriction on size,
and by reusing the same offset in
different <span class="bold"><strong>define</strong></span> statements,
overlapping varnodes are allowed. This is most often used to give
registers their standard names but could be used to label any semantic
variable that might need to be accessed globally by the
processor. Overlapping register sequences like the x86 EAX/AX/AL can
be easily modeled with overlapping varnode definitions.
</p>
<p>
Here is a typical example of register definition:
</p>
<div class="informalexample"><pre class="programlisting">
define register offset=0 size=4
[EAX ECX EDX EBX ESP EBP ESI EDI ];
define register offset=0 size=2
[AX _ CX _ DX _ BX _ SP _ BP _ SI _ DI];
define register offset=0 size=1
[AL AH _ _ CL CH _ _ DL DH _ _ BL BH ];
</pre></div>
<p>
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875464736"></a>4.5. Bit Range Registers</h3></div></div></div>
<p>
Many processors define registers that either consist of a single bit
or otherwise don't use an integral number of bytes. A recurring
example in many processors is the status register which is further
subdivided into the overflow and result flags for the arithmetic
instructions. These flags are typically have labels like ZF for the
zero flag or CF for the carry flag and can be considered logical
registers contained within the status register. SLEIGH allows
registers to be defined like this using
the <span class="bold"><strong>define bitrange</strong></span> statement, but
there are some important caveats with its use. A bit register like
this is problematic for the underlying p-code instructions that SLEIGH
models because the smallest object they can manipulate directly is a
byte. In order to manipulate single bits, p-code must use a
combination of bitwise logical, extension, and truncation
operations. So a register defined as a bit range is not really a
varnode as described in <a class="xref" href="sleigh.html#sleigh_varnodes" title="1.2. Varnodes">Section 1.2, &#8220;Varnodes&#8221;</a>, but is
really just a signal to the SLEIGH compiler to fill in the proper
operators to simulate the bit manipulation. Using this feature may
greatly increase the complexity of the compiled specification with
little indication within the specification file itself.
</p>
<div class="informalexample"><pre class="programlisting">
define register offset=0x180 size=4 [ statusreg ];
define bitrange zf=statusreg[10,1]
cf=statusreg[11,1]
sf=statusreg[12,1];
</pre></div>
<p>
</p>
<p>
A bit range register must be defined on top of another normal
register. In this example, <span class="emphasis"><em>statusreg</em></span> is defined
first as a 4 byte register, and the bit registers themselves are built
by the following <span class="bold"><strong>define bitrange</strong></span>
statement. A single bit register definition consists of an identifier
for the register, followed by &#8216;=&#8217;, then the name of the register
containing the bits, and finally a pair of numbers in square
brackets. The first number indicates the lowest significant bit in the
containing register of the bit range, where bit 0 is the least
significant bit. The second number indicates the number of bits in the
new register. Multiple definitions can be included in a
single <span class="bold"><strong>define bitrange</strong></span> statement, and
the command is finally terminated with a semicolon. In the example,
three new registers are defined on top
of <span class="emphasis"><em>statusreg</em></span>, each made up of 1 bit. The new
registers <span class="emphasis"><em>zf</em></span>, <span class="emphasis"><em>cf</em></span>,
and <span class="emphasis"><em>sf</em></span> represent the tenth, eleventh, and twelfth
bit of <span class="emphasis"><em>statusreg</em></span> respectively.
</p>
<p>
The syntax for defining a new bit register is consistent with the
pseudo bit range operator, described in
<a class="xref" href="sleigh_constructors.html#sleigh_bitrange_operator" title="7.7.1.5. Bit Range Operator">Section 7.7.1.5, &#8220;Bit Range Operator&#8221;</a>, and the resulting symbol
is really just a placeholder for this operator. Whenever SLEIGH sees
this symbol it generates p-code precisely as if the designer had used
the bit range operator
instead. <a class="xref" href="sleigh_constructors.html#sleigh_bitrange_operator" title="7.7.1.5. Bit Range Operator">Section 7.7.1.5, &#8220;Bit Range Operator&#8221;</a>, provides some
additional details about how p-code is generated, which apply to the
use of bit range registers.
</p>
<p>
If a defined bit range happens to fall on byte boundaries, the new
symbol will in fact be a normal varnode, so
the <span class="bold"><strong>define bitrange</strong></span> statement can be
used as an alternate syntax for defining overlapping registers.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875451744"></a>4.6. User-Defined Operations</h3></div></div></div>
<p>
The specification designer can define new p-code operations using
a <span class="bold"><strong>define pcodeop</strong></span> statement. This
statement automatically reserves an internal form for the new p-code
operation and associates an identifier with it. This identifier can
then be used in semantic expressions (see
<a class="xref" href="sleigh_constructors.html#sleigh_userdef_op" title="7.7.1.8. User-Defined Operations">Section 7.7.1.8, &#8220;User-Defined Operations&#8221;</a>). The following example defines a
new p-code operation <span class="emphasis"><em>arctan</em></span>.
</p>
<div class="informalexample"><pre class="programlisting">
define pcodeop arctan;
</pre></div>
<p>
</p>
<p>
This construction should be used sparingly. The definition does not
specify how the new operation is supposed to actually manipulate data,
and any analysis routines cannot know what the specification designer
intended. The operation will be treated as a black box. It will hold
its place in syntax trees, and the routines will understand how data
flows into and out of it. But, no other analysis will be possible.
</p>
<p>
New operations should be defined only after considering the above
points and the general philosophy of p-code. The designer should have
a detailed description of the new operation in mind, even though this
cannot be put in the specification. If it all possible, the operation
should be atomic, with specific inputs and outputs, and with no
side-effects. The most common use of a new operation is to encapsulate
actions that are too esoteric or too complicated to implement.
</p>
</div>
</div>
<div class="navfooter">
<hr>
<table width="100%" summary="Navigation footer">
<tr>
<td width="40%" align="left">
<a accesskey="p" href="sleigh_preprocessing.html">Prev</a> </td>
<td width="20%" align="center"> </td>
<td width="40%" align="right"> <a accesskey="n" href="sleigh_symbols.html">Next</a>
</td>
</tr>
<tr>
<td width="40%" align="left" valign="top">3. Preprocessing </td>
<td width="20%" align="center"><a accesskey="h" href="sleigh.html">Home</a></td>
<td width="40%" align="right" valign="top"> 5. Introduction to Symbols</td>
</tr>
</table>
</div>
</body>
</html>

View File

@@ -0,0 +1,122 @@
<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=ISO-8859-1">
<title>2. Basic Specification Layout</title>
<link rel="stylesheet" type="text/css" href="Frontpage.css">
<link rel="stylesheet" type="text/css" href="languages.css">
<meta name="generator" content="DocBook XSL Stylesheets V1.78.1">
<link rel="home" href="sleigh.html" title="SLEIGH">
<link rel="up" href="sleigh.html" title="SLEIGH">
<link rel="prev" href="sleigh.html" title="SLEIGH">
<link rel="next" href="sleigh_preprocessing.html" title="3. Preprocessing">
</head>
<body bgcolor="white" text="black" link="#0000FF" vlink="#840084" alink="#0000FF">
<div class="navheader">
<table width="100%" summary="Navigation header">
<tr><th colspan="3" align="center">2. Basic Specification Layout</th></tr>
<tr>
<td width="20%" align="left">
<a accesskey="p" href="sleigh.html">Prev</a> </td>
<th width="60%" align="center"> </th>
<td width="20%" align="right"> <a accesskey="n" href="sleigh_preprocessing.html">Next</a>
</td>
</tr>
</table>
<hr>
</div>
<div class="sect1">
<div class="titlepage"><div><div><h2 class="title" style="clear: both">
<a name="sleigh_layout"></a>2. Basic Specification Layout</h2></div></div></div>
<p>
A SLEIGH specification is typically contained in a single file,
although see <a class="xref" href="sleigh_preprocessing.html#sleigh_including_files" title="3.1. Including Files">Section 3.1, &#8220;Including Files&#8221;</a>. The file must
follow a specific format as parsed by the SLEIGH compiler. In this
section, we list the basic formatting rules for this file as enforced
by the compiler.
</p>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875562464"></a>2.1. Comments</h3></div></div></div>
<p>
Comments start with the &#8216;#&#8217; character and continue to the end of the
line. Comments can appear anywhere except the <span class="emphasis"><em>display section</em></span> of a
constructor (see <a class="xref" href="sleigh_constructors.html#sleigh_display_section" title="7.3. The Display Section">Section 7.3, &#8220;The Display Section&#8221;</a>) where the &#8216;#&#8217; character will be
interpreted as something that should be printed in disassembly.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875560064"></a>2.2. Identifiers</h3></div></div></div>
<p>
Identifiers are made up of letters a-z, capitals A-Z, digits 0-9 and
the characters &#8216;.&#8217; and &#8216;_&#8217;. An identifier can use these characters in
any order and for any length, but it must not start with a digit.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875558464"></a>2.3. Strings</h3></div></div></div>
<p>
String literals can be used, when specifying names and when specifying
how disassembly should be printed, so that special characters are
treated as literals. Strings are surrounded by the double quote
character &#8216;&#8221;&#8217; and all characters in between lose their special
meaning.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875556736"></a>2.4. Integers</h3></div></div></div>
<p>
Integers are specified either in a decimal format or in a standard
<span class="emphasis"><em>C-style</em></span> hexadecimal format by prepending the
number with &#8220;0x&#8221;. Alternately, a binary representation of an integer
can be given by prepending the string of '0' and '1' characters with "0b".
</p>
<div class="informalexample"><pre class="programlisting">
1006789
0xF5CC5
0xf5cc5
0b11110101110011000101
</pre></div>
<p>
Numbers are treated as unsigned
except when used in patterns where they are treated as signed (see
<a class="xref" href="sleigh_constructors.html#sleigh_bit_pattern" title="7.4. The Bit Pattern Section">Section 7.4, &#8220;The Bit Pattern Section&#8221;</a>). The number of bytes used to
encode the integer when specifying the semantics of an instruction is
inferred from other parts of the syntax (see
<a class="xref" href="sleigh_constructors.html#sleigh_display_section" title="7.3. The Display Section">Section 7.3, &#8220;The Display Section&#8221;</a>). Otherwise, integers should
be thought of as having arbitrary precision. Currently, SLEIGH stores
integers internally with 64 bits of precision.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875552544"></a>2.5. White Space</h3></div></div></div>
<p>
White space characters include space, tab, line-feed, vertical
line-feed, and carriage-return (&#8216; &#8216;, &#8216;\t&#8217;, &#8216;\r&#8217;, &#8216;\v&#8217;,
&#8216;\n&#8217;). Variations in spacing have no effect on the parsing of the file
except in string literals.
</p>
</div>
</div>
<div class="navfooter">
<hr>
<table width="100%" summary="Navigation footer">
<tr>
<td width="40%" align="left">
<a accesskey="p" href="sleigh.html">Prev</a> </td>
<td width="20%" align="center"> </td>
<td width="40%" align="right"> <a accesskey="n" href="sleigh_preprocessing.html">Next</a>
</td>
</tr>
<tr>
<td width="40%" align="left" valign="top">SLEIGH </td>
<td width="20%" align="center"><a accesskey="h" href="sleigh.html">Home</a></td>
<td width="40%" align="right" valign="top"> 3. Preprocessing</td>
</tr>
</table>
</div>
</body>
</html>

View File

@@ -0,0 +1,214 @@
<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=ISO-8859-1">
<title>3. Preprocessing</title>
<link rel="stylesheet" type="text/css" href="Frontpage.css">
<link rel="stylesheet" type="text/css" href="languages.css">
<meta name="generator" content="DocBook XSL Stylesheets V1.78.1">
<link rel="home" href="sleigh.html" title="SLEIGH">
<link rel="up" href="sleigh.html" title="SLEIGH">
<link rel="prev" href="sleigh_layout.html" title="2. Basic Specification Layout">
<link rel="next" href="sleigh_definitions.html" title="4. Basic Definitions">
</head>
<body bgcolor="white" text="black" link="#0000FF" vlink="#840084" alink="#0000FF">
<div class="navheader">
<table width="100%" summary="Navigation header">
<tr><th colspan="3" align="center">3. Preprocessing</th></tr>
<tr>
<td width="20%" align="left">
<a accesskey="p" href="sleigh_layout.html">Prev</a> </td>
<th width="60%" align="center"> </th>
<td width="20%" align="right"> <a accesskey="n" href="sleigh_definitions.html">Next</a>
</td>
</tr>
</table>
<hr>
</div>
<div class="sect1">
<div class="titlepage"><div><div><h2 class="title" style="clear: both">
<a name="sleigh_preprocessing"></a>3. Preprocessing</h2></div></div></div>
<p>
SLEIGH provides support for simple file inclusion, macros, and other
basic preprocessing functions. These are all invoked with directives
that start with the &#8216;@&#8217; character, which must be the first character
in the line.
</p>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="sleigh_including_files"></a>3.1. Including Files</h3></div></div></div>
<p>
In general a single SLEIGH specification is contained in a single
file, and the compiler is invoked on one file at a time. Multiple
files can be put together for one specification by using
the <span class="bold"><strong>@include</strong></span> directive. This must
appear at the beginning of the line and is followed by the path name
of the file to be included, enclosed in double quotes.
</p>
<div class="informalexample"><code class="code">@include "example.slaspec"</code></div>
<p>
Parsing proceeds as if the entire line is replaced with the contents
of the indicated file. Multiple inclusions are possible, and the
included files can have their
own <span class="bold"><strong>@include</strong></span> directives.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875545072"></a>3.2. Preprocessor Macros</h3></div></div></div>
<p>
SLEIGH allows simple (unparameterized) macro definitions and
expansions. A macro definition occurs on one line and starts with
the <span class="bold"><strong>@define</strong></span> directive. This is
followed by an identifier for the macro and then a string to which the
macro should expand. The string must either be a proper identifier
itself or surrounded with double quotes. The macro can then be
expanded with typical &#8220;$(identifier)&#8221; syntax at any other point in the
specification following the definition.
</p>
<div class="informalexample"><pre class="programlisting">
@define ENDIAN "big"
<span class="weak">...</span>
define endian=$(ENDIAN);
</pre></div>
<p>
This example defines a macro identified as <span class="emphasis"><em>ENDIAN</em></span>
with the string &#8220;big&#8221;, and then expands the macro in a later SLEIGH
statement. Macro definitions can also be made from the command line
and in the &#8220;.spec&#8221; file, allowing multiple specification variations to
be derived from one file. SLEIGH also has
an <span class="bold"><strong>@undef</strong></span> directive which removes the
definition of a macro from that point on in the file.
</p>
<div class="informalexample"><code class="code">@undef ENDIAN</code></div>
<p>
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875538656"></a>3.3. Conditional Compilation</h3></div></div></div>
<p>
SLEIGH supports several directives that allow conditional inclusion of
parts of a specification, based on the existence of a macro, or its
value. The lines of the specification to be conditionally included are
bounded by one of the <span class="bold"><strong>@if...</strong></span>
directives described below and at the bottom by
the <span class="bold"><strong>@endif</strong></span> directive. If the
condition described by the <span class="bold"><strong>@if...</strong></span>
directive is true, the bounded lines are evaluated as part of the
specification, otherwise they are skipped. Nesting of these directives
is allowed: a
second <span class="bold"><strong>@if...</strong></span> <span class="bold"><strong>@endif</strong></span>
pair can occur inside an initial <span class="bold"><strong>@if</strong></span>
and <span class="bold"><strong>@endif</strong></span>.
</p>
<div class="sect3">
<div class="titlepage"><div><div><h4 class="title">
<a name="idm140310875532832"></a>3.3.1. @ifdef and @ifndef</h4></div></div></div>
<p>
The <span class="bold"><strong>@ifdef</strong></span> directive is followed by a
macro identifier and evaluates to true if the macro is defined.
The <span class="bold"><strong>@ifndef</strong></span> directive is similar
except it evaluates to true if the macro identifier
is <span class="emphasis"><em>not</em></span> defined.
</p>
<div class="informalexample"><pre class="programlisting">
@ifdef ENDIAN
define endian=$(ENDIAN);
@else
define endian=little;
@endif
</pre></div>
<p>
This directive can only take a single identifier as an argument, any
other form is flagged as an error. For logically combining a test of
whether a macro is defined with other tests, use
the <span class="bold"><strong>defined</strong></span> operator in
an <span class="bold"><strong>@if</strong></span>
or <span class="bold"><strong>@elif</strong></span> directive (See below).
</p>
</div>
<div class="sect3">
<div class="titlepage"><div><div><h4 class="title">
<a name="idm140310875526896"></a>3.3.2. @if</h4></div></div></div>
<p>
The <span class="bold"><strong>@if</strong></span> directive is followed by a
boolean expression with macros as the variables and strings as the
constants. Comparisons between macros and strings are currently
limited to string equality or inequality. But individual comparisons
can be combined arbitrarily using parentheses and the boolean
operators &#8216;&amp;&amp;&#8217;, &#8216;||&#8217;, and &#8216;^^&#8217;. These represent a <span class="emphasis"><em>logical
and</em></span>, a <span class="emphasis"><em>logical or</em></span>, and
a <span class="emphasis"><em>logical exclusive-or</em></span> operation respectively. It
is possible to test whether a particular macro is defined within the
boolean expression for an <span class="bold"><strong>@if</strong></span>
directive, by using the <span class="bold"><strong>defined</strong></span>
operator. This exists as a keyword and a functional operator only
within a preprocessor boolean
expression. The <span class="bold"><strong>defined</strong></span> keyword takes
as argument a macro identifier, and it evaluates to true if the macro
is defined.
</p>
<div class="informalexample"><pre class="programlisting">
@if defined(X_EXTENSION) || (VERSION == "5")
...
@endif
</pre></div>
<p>
</p>
</div>
<div class="sect3">
<div class="titlepage"><div><div><h4 class="title">
<a name="idm140310875520240"></a>3.3.3. @else and @elif</h4></div></div></div>
<p>
An <span class="bold"><strong>@else</strong></span> directive splits the lines
bounded by an <span class="bold"><strong>@if</strong></span> directive and
an <span class="bold"><strong>@endif</strong></span> directive into two
parts. The first part is included in the processing if the
initial <span class="bold"><strong>@if</strong></span> directive evaluates to
true, otherwise the second part is included.
</p>
<p>
The <span class="bold"><strong>@elif</strong></span> directive splits the
bounded lines up as with <span class="bold"><strong>@else</strong></span>, but
the second part is included only if the
previous <span class="bold"><strong>@if</strong></span> was false and the
condition specified in the <span class="bold"><strong>@elif</strong></span>
itself is true. Between one <span class="bold"><strong>@if</strong></span>
and <span class="bold"><strong>@endif</strong></span> pair, there can be
multiple <span class="bold"><strong>@elif</strong></span> directives, but only
one <span class="bold"><strong>@else</strong></span>, which must occur after all
the <span class="bold"><strong>@elif</strong></span> directives.
</p>
<div class="informalexample"><pre class="programlisting">
@if PROCESSOR == &#8220;mips&#8221;
@ define ENDIAN &#8220;big&#8221;
@elif ((PROCESSOR==&#8221;x86&#8221;)&amp;&amp;(OS!=&#8221;win&#8221;))
@ define ENDIAN &#8220;little&#8221;
@else
@ define ENDIAN &#8220;unknown&#8221;
@endif
</pre></div>
<p>
</p>
</div>
</div>
</div>
<div class="navfooter">
<hr>
<table width="100%" summary="Navigation footer">
<tr>
<td width="40%" align="left">
<a accesskey="p" href="sleigh_layout.html">Prev</a> </td>
<td width="20%" align="center"> </td>
<td width="40%" align="right"> <a accesskey="n" href="sleigh_definitions.html">Next</a>
</td>
</tr>
<tr>
<td width="40%" align="left" valign="top">2. Basic Specification Layout </td>
<td width="20%" align="center"><a accesskey="h" href="sleigh.html">Home</a></td>
<td width="40%" align="right" valign="top"> 4. Basic Definitions</td>
</tr>
</table>
</div>
</body>
</html>

View File

@@ -0,0 +1,595 @@
<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=ISO-8859-1">
<title>9. P-code Tables</title>
<link rel="stylesheet" type="text/css" href="Frontpage.css">
<link rel="stylesheet" type="text/css" href="languages.css">
<meta name="generator" content="DocBook XSL Stylesheets V1.78.1">
<link rel="home" href="sleigh.html" title="SLEIGH">
<link rel="up" href="sleigh.html" title="SLEIGH">
<link rel="prev" href="sleigh_context.html" title="8. Using Context">
</head>
<body bgcolor="white" text="black" link="#0000FF" vlink="#840084" alink="#0000FF">
<div class="navheader">
<table width="100%" summary="Navigation header">
<tr><th colspan="3" align="center">9. P-code Tables</th></tr>
<tr>
<td width="20%" align="left">
<a accesskey="p" href="sleigh_context.html">Prev</a> </td>
<th width="60%" align="center"> </th>
<td width="20%" align="right"> </td>
</tr>
</table>
<hr>
</div>
<div class="sect1">
<div class="titlepage"><div><div><h2 class="title" style="clear: both">
<a name="sleigh_ref"></a>9. P-code Tables</h2></div></div></div>
<p>
We list all the p-code operations by name along with the syntax for
invoking them within the semantic section of a constructor definition
(see <a class="xref" href="sleigh_constructors.html#sleigh_semantic_section" title="7.7. The Semantic Section">Section 7.7, &#8220;The Semantic Section&#8221;</a>), and with a
description of the operator. The terms <span class="emphasis"><em>v0</em></span>
and <span class="emphasis"><em>v1</em></span> represent identifiers of individual input
varnodes to the operation. In terms of syntax, <span class="emphasis"><em>v0</em></span>
and <span class="emphasis"><em>v1</em></span> can be replaced with any semantic
expression, in which case the final output varnode of the expression
becomes the input to the operator. The term <span class="emphasis"><em>spc</em></span>
represents the identifier of an address space, which is a special
input to the <span class="emphasis"><em>LOAD</em></span> and <span class="emphasis"><em>STORE</em></span>
operations. The identifier of any address space can be used.
</p>
<p>
This table lists all the operators for building semantic
expressions. The operators are listed in order of precedence, highest
to lowest.
</p>
<div class="informalexample">
<div class="table">
<a name="syntaxref.htmltable"></a><p class="title"><b>Table 5. Semantic Expression Operators and Syntax</b></p>
<div class="table-contents"><table width="95%" frame="box" rules="all">
<col width="25%">
<col width="25%">
<col width="50%">
<thead><tr>
<td><span class="bold"><strong>P-code Name</strong></span></td>
<td><span class="bold"><strong>SLEIGH Syntax</strong></span></td>
<td><span class="bold"><strong>Description</strong></span></td>
</tr></thead>
<tbody>
<tr>
<td><code class="code">SUBPIECE</code></td>
<td>
<div class="informaltable">
<a name="subpieceref.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">v0:2</code></td>
</tr>
<tr>
<td><code class="code">v0(2)</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>The least significant n bytes of v0.
Truncate least significant n bytes of
v0. Most significant bytes may be
truncated depending on result size.
</td>
</tr>
<tr>
<td><code class="code">(simulated)</code></td>
<td><code class="code">v0[6,1]</code></td>
<td>Extract a range of bits from v0,
putting result in a minimum number of
bytes. The bracketed numbers give
respectively, the least significant
bit and the number of bits in the
range.
</td>
</tr>
<tr>
<td><code class="code">LOAD</code></td>
<td>
<div class="informaltable">
<a name="loadref.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">* v1</code></td>
</tr>
<tr>
<td><code class="code">*[spc]v1</code></td>
</tr>
<tr>
<td><code class="code">*:2 v1</code></td>
</tr>
<tr>
<td><code class="code">*[spc]:2 v1</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>Dereference v1 as pointer into
default space. Optionally specify
space to load from and size of data
in bytes.
</td>
</tr>
<tr>
<td><code class="code">BOOL_NEGATE</code></td>
<td><code class="code">!v0</code></td>
<td>Negation of boolean value v0.</td>
</tr>
<tr>
<td><code class="code">INT_NEGATE</code></td>
<td><code class="code">~v0</code></td>
<td>Bitwise negation of v0.</td>
</tr>
<tr>
<td><code class="code">INT_2COMP</code></td>
<td><code class="code">-v0</code></td>
<td>Twos complement of v0.</td>
</tr>
<tr>
<td><code class="code">FLOAT_NEG</code></td>
<td><code class="code">f- v0</code></td>
<td>Additive inverse of v0 as a floating-point number.</td>
</tr>
<tr>
<td><code class="code">INT_MULT</code></td>
<td><code class="code">v0 * v1</code></td>
<td>Integer multiplication of v0 and v1.</td>
</tr>
<tr>
<td><code class="code">INT_DIV</code></td>
<td><code class="code">v0 / v1</code></td>
<td>Unsigned division of v0 by v1.</td>
</tr>
<tr>
<td><code class="code">INT_SDIV</code></td>
<td><code class="code">v0 s/ v1</code></td>
<td>Signed division of v0 by v1.</td>
</tr>
<tr>
<td><code class="code">INT_REM</code></td>
<td><code class="code">v0 % v1</code></td>
<td>Unsigned remainder of v0 modulo v1.</td>
</tr>
<tr>
<td><code class="code">INT_SREM</code></td>
<td><code class="code">v0 s% v1</code></td>
<td>Signed remainder of v0 modulo v1.</td>
</tr>
<tr>
<td><code class="code">FLOAT_DIV</code></td>
<td><code class="code">v0 f/ v1</code></td>
<td>Division of v0 by v1 as floating-point numbers.</td>
</tr>
<tr>
<td><code class="code">FLOAT_MULT</code></td>
<td><code class="code">v0 f* v1</code></td>
<td>Multiplication of v0 and v1 as floating-point numbers.</td>
</tr>
<tr>
<td><code class="code">INT_ADD</code></td>
<td><code class="code">v0 + v1</code></td>
<td>Addition of v0 and v1 as integers.</td>
</tr>
<tr>
<td><code class="code">INT_SUB</code></td>
<td><code class="code">v0 - v1</code></td>
<td>Subtraction of v1 from v0 as integers.</td>
</tr>
<tr>
<td><code class="code">FLOAT_ADD</code></td>
<td><code class="code">v0 f+ v1</code></td>
<td>Addition of v0 and v1 as floating-point numbers.</td>
</tr>
<tr>
<td><code class="code">FLOAT_SUB</code></td>
<td><code class="code">v0 f- v1</code></td>
<td>Subtraction of v1 from v0 as floating-point numbers.</td>
</tr>
<tr>
<td><code class="code">INT_LEFT</code></td>
<td><code class="code">v0 &lt;&lt; v1</code></td>
<td>Left shift of v0 by v1 bits.</td>
</tr>
<tr>
<td><code class="code">INT_RIGHT</code></td>
<td><code class="code">v0 &gt;&gt; v1</code></td>
<td>Unsigned (logical) right shift of v0 by v1 bits.</td>
</tr>
<tr>
<td><code class="code">INT_SRIGHT</code></td>
<td><code class="code">v0 s&gt;&gt; v1</code></td>
<td>Signed (arithmetic) right shift of v0 by b1 bits.</td>
</tr>
<tr>
<td><code class="code">INT_SLESS</code></td>
<td>
<div class="informaltable">
<a name="slessref.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">v0 s&lt; v1</code></td>
</tr>
<tr>
<td><code class="code">v1 s&gt; v0</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>True if v0 is less than v1 as a signed integer.</td>
</tr>
<tr>
<td><code class="code">INT_SLESSEQUAL</code></td>
<td>
<div class="informaltable">
<a name="slessequalref.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">v0 s&lt;= v1</code></td>
</tr>
<tr>
<td><code class="code">v1 s&gt;= v0</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>True if v0 is less than or equal to v1 as a signed integer.</td>
</tr>
<tr>
<td><code class="code">INT_LESS</code></td>
<td>
<div class="informaltable">
<a name="lessref.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">v0 &lt; v1</code></td>
</tr>
<tr>
<td><code class="code">v1 &gt; v0</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>True if v0 is less than v1 as an unsigned integer.</td>
</tr>
<tr>
<td><code class="code">INT_LESSEQUAL</code></td>
<td>
<div class="informaltable">
<a name="lessequalref.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">v0 &lt;= v1</code></td>
</tr>
<tr>
<td><code class="code">v1 &gt;= v0</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>True if v0 is less than or equal to v1 as an unsigned integer.</td>
</tr>
<tr>
<td><code class="code">FLOAT_LESS</code></td>
<td>
<div class="informaltable">
<a name="flessref.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">v0 f&lt; v1</code></td>
</tr>
<tr>
<td><code class="code">v1 f&gt; v0</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>True if v0 is less than v1 viewed as floating-point numbers.</td>
</tr>
<tr>
<td><code class="code">FLOAT_LESSEQUAL</code></td>
<td>
<div class="informaltable">
<a name="flessequalref.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">v0 f&lt;= v1</code></td>
</tr>
<tr>
<td><code class="code">v1 f&gt;= v0</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>True if v0 is less than or equal to v1 as floating-point.</td>
</tr>
<tr>
<td><code class="code">INT_EQUAL</code></td>
<td><code class="code">v0 == v1</code></td>
<td>True if v0 equals v1.</td>
</tr>
<tr>
<td><code class="code">INT_NOTEQUAL</code></td>
<td><code class="code">v0 != v1</code></td>
<td>True if v0 does not equal v1.</td>
</tr>
<tr>
<td><code class="code">FLOAT_EQUAL</code></td>
<td><code class="code">v0 f== v1</code></td>
<td>True if v0 equals v1 viewed as floating-point numbers.</td>
</tr>
<tr>
<td><code class="code">FLOAT_NOTEQUAL</code></td>
<td><code class="code">v0 f!= v1</code></td>
<td>True if v0 does not equal v1 viewed as floating-point numbers.</td>
</tr>
<tr>
<td><code class="code">INT_AND</code></td>
<td><code class="code">v0 &amp; v1</code></td>
<td>Bitwise Logical And of v0 with v1.</td>
</tr>
<tr>
<td><code class="code">INT_XOR</code></td>
<td><code class="code">v0 ^ v1</code></td>
<td>Bitwise Exclusive Or of v0 with v1.</td>
</tr>
<tr>
<td><code class="code">INT_OR</code></td>
<td><code class="code">v0 | v1</code></td>
<td>Bitwise Logical Or of v0 with v1.</td>
</tr>
<tr>
<td><code class="code">BOOL_XOR</code></td>
<td><code class="code">v0 ^^ v1</code></td>
<td>Exclusive-Or of booleans v0 and v1.</td>
</tr>
<tr>
<td><code class="code">BOOL_AND</code></td>
<td><code class="code">v0 &amp;&amp; v1</code></td>
<td>Logical-And of booleans v0 and v1.</td>
</tr>
<tr>
<td><code class="code">BOOL_OR</code></td>
<td><code class="code">v0 || v1</code></td>
<td>Logical-Or of booleans v0 and v1.</td>
</tr>
<tr>
<td><code class="code">INT_ZEXT</code></td>
<td><code class="code">zext(v0)</code></td>
<td>Zero extension of v0.</td>
</tr>
<tr>
<td><code class="code">INT_SEXT</code></td>
<td><code class="code">sext(v0)</code></td>
<td>Sign extension of v0.</td>
</tr>
<tr>
<td><code class="code">INT_CARRY</code></td>
<td><code class="code">carry(v0,v1)</code></td>
<td>True if adding v0 and v1 would produce an unsigned carry.</td>
</tr>
<tr>
<td><code class="code">INT_SCARRY</code></td>
<td><code class="code">scarry(v0,v1)</code></td>
<td>True if adding v0 and v1 would produce a signed carry.</td>
</tr>
<tr>
<td><code class="code">INT_SBORROW</code></td>
<td><code class="code">sborrow(v0,v1)</code></td>
<td>True if subtracting v1 from v0 would produce a signed borrow.</td>
</tr>
<tr>
<td><code class="code">FLOAT_NAN</code></td>
<td><code class="code">nan(v0)</code></td>
<td>True if v0 is not a valid floating-point number (NaN).</td>
</tr>
<tr>
<td><code class="code">FLOAT_ABS</code></td>
<td><code class="code">abs(v0)</code></td>
<td>Absolute value of v0 as floating point number.</td>
</tr>
<tr>
<td><code class="code">FLOAT_SQRT</code></td>
<td><code class="code">sqrt(v0)</code></td>
<td>Square root of v0 as floating-point number.</td>
</tr>
<tr>
<td><code class="code">INT2FLOAT</code></td>
<td><code class="code">int2float(v0)</code></td>
<td>Floating-point representation of v0 viewed as an integer.</td>
</tr>
<tr>
<td><code class="code">FLOAT2FLOAT</code></td>
<td><code class="code">float2float(v0)</code></td>
<td>Copy of floating-point number v0 with more or less precision.</td>
</tr>
<tr>
<td><code class="code">TRUNC</code></td>
<td><code class="code">trunc(v0)</code></td>
<td>Signed integer obtained by truncating v0.</td>
</tr>
<tr>
<td><code class="code">FLOAT_CEIL</code></td>
<td><code class="code">ceil(v0)</code></td>
<td>Nearest integer greater than v0.</td>
</tr>
<tr>
<td><code class="code">FLOAT_FLOOR</code></td>
<td><code class="code">floor(v0)</code></td>
<td>Nearest integer less than v0.</td>
</tr>
<tr>
<td><code class="code">FLOAT_ROUND</code></td>
<td><code class="code">round(v0)</code></td>
<td>Nearest integer to v0.</td>
</tr>
<tr>
<td><code class="code">CPOOLREF</code></td>
<td><code class="code">cpool(v0,...)</code></td>
<td>Access value from the constant pool.</td>
</tr>
<tr>
<td><code class="code">NEW</code></td>
<td><code class="code">newobject(v0)</code></td>
<td>Allocate object of type described by v0.</td>
</tr>
<tr>
<td><code class="code"><span class="emphasis"><em>USER_DEFINED</em></span></code></td>
<td><code class="code"><span class="emphasis"><em>ident</em></span>(v0,...)</code></td>
<td>User defined operator <span class="emphasis"><em>ident</em></span>, with functional syntax.</td>
</tr>
</tbody>
</table></div>
</div>
<br class="table-break">
</div>
<p>
</p>
<p>
The following table lists the basic forms of a semantic statement.
</p>
<div class="informalexample">
<div class="table">
<a name="statementref.htmltable"></a><p class="title"><b>Table 6. Basic Statements and Associated Operators</b></p>
<div class="table-contents"><table width="95%" frame="box" rules="all">
<col width="25%">
<col width="25%">
<col width="50%">
<thead><tr>
<td><span class="bold"><strong>P-code Name</strong></span></td>
<td><span class="bold"><strong>SLEIGH Syntax</strong></span></td>
<td><span class="bold"><strong>Description</strong></span></td>
</tr></thead>
<tbody>
<tr>
<td><code class="code">COPY, <span class="emphasis"><em>other</em></span></code></td>
<td><code class="code">v0 = v1;</code></td>
<td>Assignment of v1 to v0.</td>
</tr>
<tr>
<td><code class="code">STORE</code></td>
<td>
<div class="informaltable">
<a name="storeref.htmltable"></a><table frame="none"><tbody>
<tr>
<td><code class="code">*v0 = v1</code></td>
</tr>
<tr>
<td><code class="code">*[spc]v0 = v1;</code></td>
</tr>
<tr>
<td><code class="code">*:4 v0 = v1;</code></td>
</tr>
<tr>
<td><code class="code">*[spc]:4 v0 = v1;</code></td>
</tr>
</tbody></table>
</div>
</td>
<td>Store v1 in default space using v0
As pointer. Optionally specify space
to store in and size of data in
bytes.
</td>
</tr>
<tr>
<td><code class="code"><span class="emphasis"><em>USER_DEFINED</em></span></code></td>
<td><code class="code"><span class="emphasis"><em>ident</em></span>(v0,...);</code></td>
<td>Invoke user-defined operation ident as a standalone statement, with no output.</td>
</tr>
<tr>
<td></td>
<td><code class="code">v0[8,1] = v1;</code></td>
<td>Fill a bit range within v0 using v1, leaving the rest of v0 unchanged.</td>
</tr>
<tr>
<td></td>
<td><code class="code"><span class="emphasis"><em>ident</em></span>(v0,...);</code></td>
<td>Invoke the macro named <span class="emphasis"><em>ident</em></span>.</td>
</tr>
<tr>
<td></td>
<td><code class="code">build <span class="emphasis"><em>ident</em></span>;</code></td>
<td>Execute the p-code to build operand <span class="emphasis"><em>ident</em></span>.</td>
</tr>
<tr>
<td></td>
<td><code class="code">delayslot(1);</code></td>
<td>Execute the p-code for the following instruction.</td>
</tr>
</tbody>
</table></div>
</div>
<br class="table-break">
</div>
<p>
</p>
<p>
The following table lists the branching operations and the statements which invoke them.
</p>
<div class="informalexample">
<div class="table">
<a name="branchref.htmltable"></a><p class="title"><b>Table 7. Branching Statements</b></p>
<div class="table-contents"><table width="95%" frame="box" rules="all">
<col width="25%">
<col width="25%">
<col width="50%">
<thead><tr>
<td><span class="bold"><strong>P-code Name</strong></span></td>
<td><span class="bold"><strong>SLEIGH Syntax</strong></span></td>
<td><span class="bold"><strong>Description</strong></span></td>
</tr></thead>
<tbody>
<tr>
<td><code class="code">BRANCH</code></td>
<td><code class="code">goto v0;</code></td>
<td>Branch execution to address of v0.</td>
</tr>
<tr>
<td><code class="code">CBRANCH</code></td>
<td><code class="code">if (v0) goto v1;</code></td>
<td>Branch execution to address of v1 if v0 equals 1 (true).</td>
</tr>
<tr>
<td><code class="code">BRANCHIND</code></td>
<td><code class="code">goto [v0];</code></td>
<td>Branch execution to v0 viewed as an offset in current space.</td>
</tr>
<tr>
<td><code class="code">CALL</code></td>
<td><code class="code">call v0;</code></td>
<td>Branch execution to address of v0. Hint that branch is subroutine call.</td>
</tr>
<tr>
<td><code class="code">CALLIND</code></td>
<td><code class="code">call [v0];</code></td>
<td>Branch execution to v0 viewed as an offset in current space. Hint that branch is subroutine call.</td>
</tr>
<tr>
<td><code class="code">RETURN</code></td>
<td><code class="code">return [v0];</code></td>
<td>Branch execution to v0 viewed as an offset in current space. Hint that branch is a subroutine return.</td>
</tr>
</tbody>
</table></div>
</div>
<br class="table-break">
</div>
<p>
</p>
</div>
<div class="navfooter">
<hr>
<table width="100%" summary="Navigation footer">
<tr>
<td width="40%" align="left">
<a accesskey="p" href="sleigh_context.html">Prev</a> </td>
<td width="20%" align="center"> </td>
<td width="40%" align="right"> </td>
</tr>
<tr>
<td width="40%" align="left" valign="top">8. Using Context </td>
<td width="20%" align="center"><a accesskey="h" href="sleigh.html">Home</a></td>
<td width="40%" align="right" valign="top"> </td>
</tr>
</table>
</div>
</body>
</html>

View File

@@ -0,0 +1,224 @@
<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=ISO-8859-1">
<title>5. Introduction to Symbols</title>
<link rel="stylesheet" type="text/css" href="Frontpage.css">
<link rel="stylesheet" type="text/css" href="languages.css">
<meta name="generator" content="DocBook XSL Stylesheets V1.78.1">
<link rel="home" href="sleigh.html" title="SLEIGH">
<link rel="up" href="sleigh.html" title="SLEIGH">
<link rel="prev" href="sleigh_definitions.html" title="4. Basic Definitions">
<link rel="next" href="sleigh_tokens.html" title="6. Tokens and Fields">
</head>
<body bgcolor="white" text="black" link="#0000FF" vlink="#840084" alink="#0000FF">
<div class="navheader">
<table width="100%" summary="Navigation header">
<tr><th colspan="3" align="center">5. Introduction to Symbols</th></tr>
<tr>
<td width="20%" align="left">
<a accesskey="p" href="sleigh_definitions.html">Prev</a> </td>
<th width="60%" align="center"> </th>
<td width="20%" align="right"> <a accesskey="n" href="sleigh_tokens.html">Next</a>
</td>
</tr>
</table>
<hr>
</div>
<div class="sect1">
<div class="titlepage"><div><div><h2 class="title" style="clear: both">
<a name="sleigh_symbols"></a>5. Introduction to Symbols</h2></div></div></div>
<p>
After the definition section, we are prepared to start writing the
body of the specification. This part of the specification shows how
the bits in an instruction break down into opcodes, operands,
immediate values, and the other pieces of an instruction. Then once
this is figured out, the specification must also describe exactly how
the processor would manipulate the data and operands if this
particular instruction were executed. All of SLEIGH revolves around
these two major tasks of disassembling and following semantics. It
should come as no surprise then that the primary symbols defined and
manipulated in the specification all have two key properties.
</p>
<div class="informalexample"><div class="orderedlist"><ol class="orderedlist compact" type="1">
<li class="listitem">
How does the symbol get displayed as part of the disassembly?
</li>
<li class="listitem">
What semantic variable is associated with the symbol, and how is it constructed?
</li>
</ol></div></div>
<p>
Formally a <span class="emphasis"><em>Specific Symbol</em></span> is defined as an identifier associated with
</p>
<div class="informalexample"><div class="orderedlist"><ol class="orderedlist compact" type="1">
<li class="listitem">
A string displayed in disassembly.
</li>
<li class="listitem">
varnode used in semantic actions, and any p-code used to construct that varnode.
</li>
</ol></div></div>
<p>
The named registers that we defined earlier are the simplest examples
of specific symbols (see
<a class="xref" href="sleigh_definitions.html#sleigh_naming_registers" title="4.4. Naming Registers">Section 4.4, &#8220;Naming Registers&#8221;</a>). The symbol identifier
itself is the string that will get printed in disassembly and the
varnode associated with the symbol is the one constructed by the
define statement.
</p>
<p>
The other crucial part of the specification is how to map from the
bits of a particular instruction to the specific symbols that
apply. To this end we have the <span class="emphasis"><em>Family Symbol</em></span>,
which is defined as an identifier associated with a map from machine
instructions to specific symbols.
</p>
<div class="informalexample">
<span class="bold"><strong>Family Symbol:</strong></span> Instruction Encodings =&gt; Specific Symbols
</div>
<p>
The set of instruction encodings that map to a single specific symbol
is called an <span class="emphasis"><em>instruction pattern</em></span> and is described
more fully in <a class="xref" href="sleigh_constructors.html#sleigh_bit_pattern" title="7.4. The Bit Pattern Section">Section 7.4, &#8220;The Bit Pattern Section&#8221;</a>. In most cases, this
can be thought of as a mask on the bits of the instruction and a value
that the remaining unmasked bits must match. At any rate, the family
symbol identifier, when taken out of context, represents the entire
collection of specific symbols involved in this map. But in the
context of a specific instruction, the identifier represents the one
specific symbol associated with the encoding of that instruction by
the family symbol map.
</p>
<p>
Given these maps, the idea of the specification is to build up more
and more complicated family symbols until we have a single root
symbol. This gives us a single map from the bits of an instruction to
the full disassembly of it and to the sequence of p-code instructions
that simulate the instruction.
</p>
<p>
The symbol responsible for combining smaller family symbols is called
a <span class="emphasis"><em>table</em></span>, which is fully described in
<a class="xref" href="sleigh_constructors.html#sleigh_tables" title="7.8. Tables">Section 7.8, &#8220;Tables&#8221;</a>. Any <span class="emphasis"><em>table</em></span> symbol
can be used in the definition of other <span class="emphasis"><em>table</em></span>
symbols until the root symbol is fully described. The root symbol has
the predefined identifier <span class="emphasis"><em>instruction</em></span>.
</p>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875423632"></a>5.1. Notes on Namespaces</h3></div></div></div>
<p>
Almost all identifiers live in the same global "scope". The global scope includes
</p>
<div class="informalexample"><div class="itemizedlist"><ul class="itemizedlist compact" style="list-style-type: bullet; ">
<li class="listitem" style="list-style-type: disc">
Names of address spaces
</li>
<li class="listitem" style="list-style-type: disc">
Names of tokens
</li>
<li class="listitem" style="list-style-type: disc">
Names of fields
</li>
<li class="listitem" style="list-style-type: disc">
Names of user-defined p-code ops
</li>
<li class="listitem" style="list-style-type: disc">
Names of registers
</li>
<li class="listitem" style="list-style-type: disc">
Names of macros (see <a class="xref" href="sleigh_constructors.html#sleigh_macros" title="7.9. P-code Macros">Section 7.9, &#8220;P-code Macros&#8221;</a>)
</li>
<li class="listitem" style="list-style-type: disc">
Names of tables (see <a class="xref" href="sleigh_constructors.html#sleigh_tables" title="7.8. Tables">Section 7.8, &#8220;Tables&#8221;</a>)
</li>
</ul></div></div>
<p>
All of the names in this scope must be unique. Each
individual <span class="emphasis"><em>constructor</em></span> (defined in <a class="xref" href="sleigh_constructors.html" title="7. Constructors">Section 7, &#8220;Constructors&#8221;</a>)
defines a local scope for operand names. As with most languages, a
local symbol with the same name as a global
symbol <span class="emphasis"><em>hides</em></span> the global symbol while that scope
is in affect.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="sleigh_predefined_symbols"></a>5.2. Predefined Symbols</h3></div></div></div>
<p>
We list all of the symbols that are predefined by SLEIGH.
</p>
<div class="informalexample">
<div class="table">
<a name="predefine.htmltable"></a><p class="title"><b>Table 2. Predefined Symbols</b></p>
<div class="table-contents"><table width="80%" frame="box" rules="all">
<col width="30%">
<col width="70%">
<thead><tr>
<td><span class="bold"><strong>Identifier</strong></span></td>
<td><span class="bold"><strong>Meaning</strong></span></td>
</tr></thead>
<tbody>
<tr>
<td><code class="code">instruction</code></td>
<td>The root instruction table.</td>
</tr>
<tr>
<td><code class="code">const</code></td>
<td>Special address space for building constant varnodes.</td>
</tr>
<tr>
<td><code class="code">unique</code></td>
<td>Address space for allocating temporary registers.</td>
</tr>
<tr>
<td><code class="code">inst_start</code></td>
<td>Offset of the address of the current instruction.</td>
</tr>
<tr>
<td><code class="code">inst_next</code></td>
<td>Offset of the address of the next instruction.</td>
</tr>
<tr>
<td><code class="code">epsilon</code></td>
<td>A special identifier indicating an empty bit pattern.</td>
</tr>
</tbody>
</table></div>
</div>
<br class="table-break">
</div>
<p>
The most important of these to be aware of
are <span class="emphasis"><em>inst_start</em></span>
and <span class="emphasis"><em>inst_next</em></span>. These are family symbols which map
in the context of particular instruction to the integer offset of
either the address of the instruction or the address of the next
instruction respectively. These are used in any relative branching
situation. The other symbols are rarely
used. The <span class="emphasis"><em>const</em></span> and <span class="emphasis"><em>unique</em></span>
identifiers are address spaces. The <span class="emphasis"><em>epsilon</em></span>
identifier is inherited from SLED and is a specific symbol equivalent
to the constant zero. The <span class="emphasis"><em>instruction</em></span> identifier
is the root instruction table.
</p>
</div>
</div>
<div class="navfooter">
<hr>
<table width="100%" summary="Navigation footer">
<tr>
<td width="40%" align="left">
<a accesskey="p" href="sleigh_definitions.html">Prev</a> </td>
<td width="20%" align="center"> </td>
<td width="40%" align="right"> <a accesskey="n" href="sleigh_tokens.html">Next</a>
</td>
</tr>
<tr>
<td width="40%" align="left" valign="top">4. Basic Definitions </td>
<td width="20%" align="center"><a accesskey="h" href="sleigh.html">Home</a></td>
<td width="40%" align="right" valign="top"> 6. Tokens and Fields</td>
</tr>
</table>
</div>
</body>
</html>

View File

@@ -0,0 +1,271 @@
<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=ISO-8859-1">
<title>6. Tokens and Fields</title>
<link rel="stylesheet" type="text/css" href="Frontpage.css">
<link rel="stylesheet" type="text/css" href="languages.css">
<meta name="generator" content="DocBook XSL Stylesheets V1.78.1">
<link rel="home" href="sleigh.html" title="SLEIGH">
<link rel="up" href="sleigh.html" title="SLEIGH">
<link rel="prev" href="sleigh_symbols.html" title="5. Introduction to Symbols">
<link rel="next" href="sleigh_constructors.html" title="7. Constructors">
</head>
<body bgcolor="white" text="black" link="#0000FF" vlink="#840084" alink="#0000FF">
<div class="navheader">
<table width="100%" summary="Navigation header">
<tr><th colspan="3" align="center">6. Tokens and Fields</th></tr>
<tr>
<td width="20%" align="left">
<a accesskey="p" href="sleigh_symbols.html">Prev</a> </td>
<th width="60%" align="center"> </th>
<td width="20%" align="right"> <a accesskey="n" href="sleigh_constructors.html">Next</a>
</td>
</tr>
</table>
<hr>
</div>
<div class="sect1">
<div class="titlepage"><div><div><h2 class="title" style="clear: both">
<a name="sleigh_tokens"></a>6. Tokens and Fields</h2></div></div></div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="sleigh_defining_tokens"></a>6.1. Defining Tokens and Fields</h3></div></div></div>
<p>
A <span class="emphasis"><em>token</em></span> is one of the byte-sized pieces that make
up the machine code instructions being modeled.
Instruction <span class="emphasis"><em>fields</em></span> must be defined on top of
them. A <span class="emphasis"><em>field</em></span> is a logical range of bits within
an instruction that can specify an opcode, or an operand etc. Together
tokens and fields determine the basic interpretation of bits and how
many bytes the instruction takes up. To define a token and the fields
associated with it, we use the <span class="bold"><strong>define
token</strong></span> statement.
</p>
<div class="informalexample"><pre class="programlisting">
define token <span class="bold"><strong>tokenname</strong></span> ( <span class="bold"><strong>integer</strong></span> )
<span class="bold"><strong>fieldname</strong></span>=(<span class="bold"><strong>integer</strong></span>,<span class="bold"><strong>integer</strong></span>) <span class="bold"><strong>attributelist</strong></span>
<span class="weak">...</span>
;
</pre></div>
<p>
</p>
<p>
The first part of the definition defines the name of a token and the
number of bits it uses (this must be a multiple of 8). Following this
there are one or more field declarations specifying the name of the
field and the range of bits within the token making up the field. The
size of a field does <span class="emphasis"><em>not</em></span> need to be a multiple of
8. The range is inclusive where the least significant bit in the token
is labeled 0. The endianess of the processor will effect this labeling
when defining tokens that are bigger than 1 byte. After each field
declaration, there can be zero or more of the following attribute
keywords:
</p>
<div class="informalexample"><pre class="programlisting">
signed
hex
dec
</pre></div>
<p>
These attributes are defined in the next section. There can be any
manner of repeats and overlaps in the fields so long as they all have
different names.
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875384864"></a>6.2. Fields as Family Symbols</h3></div></div></div>
<p>
Fields are the most basic form of family symbol; they define a natural
map from instruction bits to a specific symbol as follows. We take the
set of bits within the instruction as given by the field&#8217;s defining
range and treat them as an integer encoding. The resulting integer is
both the display portion and the semantic meaning of the specific
symbol. The display string is obtained by converting the integer into
either a decimal or hexadecimal representation (see below), and the
integer is treated as a constant varnode in any semantic action.
</p>
<p>
The attributes of the field affect the resulting specific symbol in
obvious ways. The <span class="bold"><strong>signed</strong></span> attribute
determines whether the integer encoding should be treated as just an
unsigned encoding or if a twos-complement encoding should be used to
obtain a signed integer. The <span class="bold"><strong>hex</strong></span>
or <span class="bold"><strong>dec</strong></span> attributes describe whether
the integer should be displayed with a hexadecimal or decimal
representation. The default is hexadecimal. [Currently
the <span class="bold"><strong>dec</strong></span> attribute is not supported]
</p>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="idm140310875379232"></a>6.3. Attaching Alternate Meanings to Fields</h3></div></div></div>
<p>
The default interpretation of a field is probably the most natural but
of course processors interpret fields within an instruction in a wide
variety of ways. The <span class="bold"><strong>attach</strong></span> keyword
is used to alter either the display or semantic meaning of fields into
the most common (and basic) interpretations. More complex
interpretations must be built up out of tables.
</p>
<div class="sect3">
<div class="titlepage"><div><div><h4 class="title">
<a name="idm140310875377152"></a>6.3.1. Attaching Registers</h4></div></div></div>
<p>
Probably <span class="emphasis"><em>the</em></span> most common processor interpretation
of a field is as an encoding of a particular register. In SLEIGH this
can be done with the <span class="bold"><strong>attach variables</strong></span>
statement:
</p>
<div class="informalexample"><pre class="programlisting">
attach variables <span class="bold"><strong>fieldlist registerlist</strong></span>;
</pre></div>
<p>
A <span class="emphasis"><em>fieldlist</em></span> can be a single field identifier or a
space separated list of field identifiers surrounded by square
brackets. A <span class="emphasis"><em>registerlist</em></span> must be a square bracket
surrounded and space separated list of register identifiers as created
with <span class="bold"><strong>define</strong></span> statements (see Section
<a class="xref" href="sleigh_definitions.html#sleigh_naming_registers" title="4.4. Naming Registers">Section 4.4, &#8220;Naming Registers&#8221;</a>). For each field in
the <span class="emphasis"><em>fieldlist</em></span>, instead of having the display and
semantic meaning of an integer, the field becomes a look-up table for
the given list of registers. The original integer interpretation is
used as the index into the list starting at zero, so a specific
instruction that has all the bits in the field equal to zero yields
the first register (a specific varnode) from the list as the meaning
of the field in the context of that instruction. Note that both the
display and semantic meaning of the field are now taken from the new
register.
</p>
<p>
A particular integer can remain unspecified by putting a &#8216;_&#8217; character
in the appropriate position of the register list or also if the length
of the register list is less than the integer. A specific integer
encoding of the field that is unspecified like this
does <span class="emphasis"><em>not</em></span> revert to the original semantic and
display meaning. Instead this encoding is flagged as an invalid form
of the instruction.
</p>
</div>
<div class="sect3">
<div class="titlepage"><div><div><h4 class="title">
<a name="idm140310875368784"></a>6.3.2. Attaching Other Integers</h4></div></div></div>
<p>
Sometimes a processor interprets a field as an integer but not the
integer given by the default interpretation. A different integer
interpretation of the field can be specified with
an <span class="bold"><strong>attach values</strong></span> statement.
</p>
<div class="informalexample"><pre class="programlisting">
attach values <span class="bold"><strong>fieldlist integerlist</strong></span>;
</pre></div>
<p>
The <span class="emphasis"><em>integerlist</em></span> is surrounded by square brackets
and is a space separated list of integers. In the same way that a new
register interpretation is assigned to fields with
an <span class="bold"><strong>attach variables</strong></span> statement, the
integers in the list are assigned to each field specified in
the <span class="emphasis"><em>fieldlist</em></span>. [Currently SLEIGH does not support
unspecified positions in the list using a &#8216;_&#8217;]
</p>
</div>
<div class="sect3">
<div class="titlepage"><div><div><h4 class="title">
<a name="idm140310875363504"></a>6.3.3. Attaching Names</h4></div></div></div>
<p>
It is possible to just modify the display characteristics of a field
without changing the semantic meaning. The need for this is rare, but
it is possible to treat a field as having influence on the display of
the disassembly but having no influence on the semantics. Even if the
bits of the field do have some semantic meaning, sometimes it is
appropriate to define overlapping fields, one of which is defined to
have no semantic meaning. The most convenient way to break down the
required disassembly may not be the most convenient way to break down
the semantics. It is also possible to have symbols with semantic
meaning but no display meaning (see <a class="xref" href="sleigh_constructors.html#sleigh_invisible_operands" title="7.4.5. Invisible Operands">Section 7.4.5, &#8220;Invisible Operands&#8221;</a>).
</p>
<p>
At any rate we can list the display interpretation of a field directly
with an <span class="bold"><strong>attach names</strong></span> statement.
</p>
<div class="informalexample"><pre class="programlisting">
attach names <span class="bold"><strong>fieldlist stringlist</strong></span>;
</pre></div>
<p>
The <span class="emphasis"><em>stringlist</em></span> is assigned to each of the fields
in the same manner as the <span class="bold"><strong>attach
variables</strong></span> and <span class="bold"><strong>attach
values</strong></span> statements. A specific encoding of the field now
displays as the string in the list at that integer position. Field
values greater than the size of the list are interpreted as invalid
encodings.
</p>
</div>
</div>
<div class="sect2">
<div class="titlepage"><div><div><h3 class="title">
<a name="sleigh_context_variables"></a>6.4. Context Variables</h3></div></div></div>
<p>
SLEIGH supports the concept of <span class="emphasis"><em>context
variables</em></span>. For the most part processor instructions can be
unambiguously decoded by examining only the bits of the instruction
encoding. But in some cases, decoding may depend on the state of
processor. Typically, the processor will have some set of status flags
that indicate what mode is being used to process instructions. In
terms of SLEIGH, a context variable is a <span class="emphasis"><em>field</em></span>
which is defined on top of a register rather than the instruction
encoding (token).
</p>
<div class="informalexample"><pre class="programlisting">
define context <span class="bold"><strong>contextreg</strong></span>
<span class="bold"><strong>fieldname</strong></span>=(<span class="bold"><strong>integer</strong></span>,<span class="bold"><strong>integer</strong></span>) <span class="bold"><strong>attributelist</strong></span>
<span class="weak">...</span>
;
</pre></div>
<p>
</p>
<p>
Context variables are defined with a <span class="bold"><strong>define
context</strong></span> statement. The keywords must be followed by the
name of a defined register. The remaining part of the definition is
nearly identical to the normal definition of fields. Each context
variable defined on this register is listed in turn, specifying the
name, the bit range, and any attributes. All the normal field attributes,
<span class="bold"><strong>signed</strong></span>, <span class="bold"><strong>dec</strong></span>, and
<span class="bold"><strong>hex</strong></span>, can also be used for context variables.
</p>
<p>
Context variables introduce a new, dedicated, attribute: <span class="bold"><strong>noflow</strong></span>.
By default, globally setting a context variable affects instruction decoding
from the point of the change, forward,
following the flow of the instructions, but if the variable is labeled as
<span class="bold"><strong>noflow</strong></span>, any change is limited to a
single instruction. (See <a class="xref" href="sleigh_context.html#sleigh_contextflow" title="8.3.1. Context Flow">Section 8.3.1, &#8220;Context Flow&#8221;</a>)
</p>
<p>
Once the context variable is defined, in terms of the specification
syntax, it can be treated as if it were just another field. See
<a class="xref" href="sleigh_context.html" title="8. Using Context">Section 8, &#8220;Using Context&#8221;</a>, for a complete discussion of how to
use context variables.
</p>
</div>
</div>
<div class="navfooter">
<hr>
<table width="100%" summary="Navigation footer">
<tr>
<td width="40%" align="left">
<a accesskey="p" href="sleigh_symbols.html">Prev</a> </td>
<td width="20%" align="center"> </td>
<td width="40%" align="right"> <a accesskey="n" href="sleigh_constructors.html">Next</a>
</td>
</tr>
<tr>
<td width="40%" align="left" valign="top">5. Introduction to Symbols </td>
<td width="20%" align="center"><a accesskey="h" href="sleigh.html">Home</a></td>
<td width="40%" align="right" valign="top"> 7. Constructors</td>
</tr>
</table>
</div>
</body>
</html>