Same class of bug as the earlier min/minu fix: the displayed source
operand order contradicts the instruction syntax recorded in each
constructor's own doc comment (and LLVM/the PRM).
vminb/vminh/vminub/vminuh/vminuw/vminw, vmaxb/vmaxh/vmaxub/vmaxuh/
vmaxuw/vmaxw and shuffob/shuffoh are documented as "f ( Rtt32 , Rss32 )"
but displayed Rss before Rtt.
vaddw and vaddw:sat are documented as "vaddw ( Rss32 , Rtt32 )" but
displayed Rtt before Rss.
Includes the V62 "Rdd32,Pe4 = vminub ( Rtt32 , Rss32 )" form.
The pcodeop / inline arguments are swapped to match, so the emitted call
argument order still follows the displayed operand order. For the min,
max and add forms this is semantically neutral (all commutative); it
matters for vminubPred, whose result is defined relative to the first
operand, and for the opaque shuffob/shuffoh, whose operand roles are
positional.
Verified: SleighCompile clean, all 60 Hexagon unit tests pass. No test
expectations change (the one test pinning vminub uses R1R0 for both
sources).
Records the decode coverage this branch adds (V66 ZReg matrix extension
and V81 valign4, both decode-only stubs) and narrows the shift-amount
KNOWN ISSUE to what is actually fixed: register-form shifts now model
the 7-bit signed shift amount, other forms may still assume positive.
Two new JUnit tests pin the disassembly text emitted for representative
encodings of the recently added decoder stubs (HexagonStubCoverageTest)
and packet-context decoding (HexagonPacketCoverageTest), so a future edit
that accidentally drops or shadows a constructor will surface as a
deliberate review point rather than a silent regression.
Adds a decode-only stub for V6_valign4.
The six veqhf/veqsf qfloat-accumulating predicate forms that originally
accompanied this are not added: upstream's hexagon_hvx2.sinc implements
all of them with real p-code.
Replace the open-ended op21/op1617 patterns on the existing vassign and
Vsf/Vhf-from-qf32/qf16 constructors with the full op2123/op1620 fields,
matching the LLVM encodings exactly so neighbouring stubs don't decode
ambiguously.
LLVM prints min as "min(Rt, Rs)" not "min(Rs, Rt)". Reorder the
display tokens for the four scalar/pair min and minu constructors;
pcode already evaluates min(rt, rs).
ictagw has no destination register; Rd5 was listed in the constructor's
operand expression but never used. Remove it so the constructor only
binds the operands it actually consumes.
LLVM disassembles the negated cmp constructors as "!cmp.eq Pd, Rs, Rt"
not "cmp.eq! Pd, Rs, Rt". Quote the bang into the head token so the
display matches; pcode is unchanged.
The rol* constructors were emitting (rs << N) | zext(rs s< 0) which only
folds the sign bit into bit 0 instead of rotating the high bits back into
the low. Use the standard (x << N) | (x >> (W - N)) pattern so the
single-shot, accumulating, and bitwise variants all match the ISA.