luv

Workshop wiki

physics-simd.lisp

luvcraft/physics-simd.lisp

system luvcraft/core · 3 definitions · on GitHub

The four-wide contact kernels: what the colouring bought (#G3W7KD).

PHYSICS.LISP solves one contact at a time. Within a colour no two contacts share an awake body, so four contacts of a colour can be gathered into f32.4 lanes, solved together with the same arithmetic in the same order, and scattered back without any lane's write landing on another lane's read. That is exactly Box3D's arrangement (#T3C8FV): the constraint columns are already the struct-of-arrays the lanes want, and only the body state needs gathering.

The kernels here are generated once per instruction family from one template, so the NEON family on arm64 and the SSE family on x86-64 are the same code lowered through sb-simd's differently named packages; a scalar family always exists as the reference (#E8M3JC). A wide kernel runs the lane groups it can and hands its tail of fewer than four contacts back to the scalar kernel. Branches become blends: where the scalar kernel skips work, the wide kernel does the work into a lane it then ignores, which is why the arithmetic never divides by a value it has not first pushed away from zero.

The claim these kernels make is testable and tested: a world stepped with the wide family and the same world stepped with the scalar family reach bit-identical state (PHYSICS-TESTS.LISP). Neither path uses a fused multiply-add or an approximate reciprocal, for that reason. #7PAQ3M

in-package#:luvcraft
defmacrodefine-wide-physics-kernels
familypackage&keyblend

Define the PHYSICS-* kernel methods for family using package's f32.4 operations. blend is :bit-select where the package has F32.4-BIT-SELECT and :and-or where SSE2's U32.4 comparison masks select the float lane bits.

flet
sym
name
internnamepackage
let
f32.4
sym"F32.4"
make
sym"MAKE-F32.4"
values4
sym"F32.4-VALUES"
aref4
sym"F32.4-AREF"
add
sym"F32.4+"
sub
sym"F32.4-"
mul
sym"F32.4*"
div
sym"F32.4/"
wmax
sym"F32.4-MAX"
gt
sym"F32.4>"
lt
sym"F32.4<"
blend-form
ecaseblend
:bit-select
let
select
sym"F32.4-BIT-SELECT"
lambda
maskthenelse
`
,select,mask,then,else
:and-or
let
wand
sym"U32.4-AND"
wandc1
sym"U32.4-ANDC1"
wor
sym"U32.4-OR"
cast
sym"U32.4!"
make
sym"MAKE-F32.4"
lambda
maskthenelse
let
m
gensym"MASK"
bits
gensym"BITS"
m0
gensym"M0"
m1
gensym"M1"
m2
gensym"M2"
m3
gensym"M3"

SSE2 comparisons produce U32.4 masks, while SBCL's SSE F32.4! cast cannot reinterpret a P128 value. Do the bit selection in the integer view, then retag its four lanes as floats. In a compiled kernel the casts into U32.4 are register moves; only this final retag is expanded lane by lane.

`
let*
,m,mask
,bits
,wor
,wand,m
,cast,then
,wandc1,m
,cast,else
multiple-value-bind
,m0,m1,m2,m3
sb-ext:%simd-pack-singles,bits
,make,m0,m1,m2,m3
`
macrolet
%wide-physics-warm-start
&restargs
list*',
intern
formatnil"%~A-PHYSICS-WARM-START"family
args
%wide-physics-solve-contacts
&restargs
list*',
intern
formatnil"%~A-PHYSICS-SOLVE-CONTACTS"family
args
%wide-physics-apply-restitution
&restargs
list*',
intern
formatnil"%~A-PHYSICS-APPLY-RESTITUTION"family
args
w+
ab
list',addab
w-
ab
list',subab
w*
ab
list',mulab
w/
ab
list',divab
wmax
ab
list',wmaxab

sb-simd 2.6.7's NEON F32.4-SQRT is mis-encoded and

raises SIGILL, so the two square roots a lane group

needs go through the scalar unit lane by lane. The

result is what the vector instruction would give.

wsqrt
a
let
v
gensym
l0
gensym
l1
gensym
l2
gensym
l3
gensym
`
let
,v,a
multiple-value-bind
,l0,l1,l2,l3
,',values4,v
,',make
sqrt
the
single-float0f0
,l0
sqrt
the
single-float0f0
,l1
sqrt
the
single-float0f0
,l2
sqrt
the
single-float0f0
,l3
w>
ab
list',gtab
w<
ab
list',ltab
wblend
maskthenelse
funcall,blend-formmaskthenelse
wsplat
x
list',f32.4x
wload
arrayindex
list',aref4arrayindex
wstore
arrayindexvalue
list'setf
list',aref4arrayindex
value
wgather
arrayi0i1i2i3
list',make
list'arefarrayi0
list'arefarrayi1
list'arefarrayi2
list'arefarrayi3
wscatter
arrayi0i1i2i3value
let
a
gensym
b
gensym
c
gensym
d
gensym
`
multiple-value-bind
,a,b,c,d
,',values4,value
setf
aref,array,i0
,a
aref,array,i1
,b
aref,array,i2
,c
aref,array,i3
,d

The relative velocity at the contact point of each side.

with-relative-velocities
vraxvrayvrazvrbxvrbyvrbz
&bodybody
`
let*
,vrax
w+vax
w-
w*wayraz
w*wazray
,vray
w+vay
w-
w*wazrax
w*waxraz
,vraz
w+vaz
w-
w*waxray
w*wayrax
,vrbx
w+
w+vbxkvx
w-
w*wbyrbz
w*wbzrby
,vrby
w+
w+vbykvy
w-
w*wbzrbx
w*wbxrbz
,vrbz
w+
w+vbzkvz
w-
w*wbxrby
w*wbyrbx
,@body

Apply the impulse P to both sides' linear and angular state.

apply-impulse
pxpypz
`
setfvax
w-vax
w*ima,px
vay
w-vay
w*ima,py
vaz
w-vaz
w*ima,pz
vbx
w+vbx
w*imb,px
vby
w+vby
w*imb,py
vbz
w+vbz
w*imb,pz
wax
w-wax
w*iia
w-
w*ray,pz
w*raz,py
way
w-way
w*iia
w-
w*raz,px
w*rax,pz
waz
w-waz
w*iia
w-
w*rax,py
w*ray,px
wbx
w+wbx
w*iib
w-
w*rby,pz
w*rbz,py
wby
w+wby
w*iib
w-
w*rbz,px
w*rbx,pz
wbz
w+wbz
w*iib
w-
w*rbx,py
w*rby,px
defmethodphysics-integrate-velocities
kernels
eql,family
awakeh
declare
single-floath
records:with-columnar-buffer-storage
countrow
vxsvx
vysvy
vzsvz
wxswx
wyswy
wzswz
dampingsdamping
inverse-massesinverse-mass
awakephysics-body-columns
declare
ignorerow
optimize
speed3
safety0
let*
gravity
*h
coerce*physics-gravity*'single-float
wide-end
*4
floorcount4
one
wsplat1f0
hs
wsplath
gs
wsplatgravity
zero
wsplat0f0
declare
single-floatgravity
fixnumwide-end
loopforifixnumfrom0belowwide-endby4do
let*
scale
w/one
w+one
w*hs
wloaddampingsi
dynamic
w>
wloadinverse-massesi
zero
fall
wblenddynamicgszero
wstorevxsi
w*scale
wloadvxsi
wstorevysi
w+
w*scale
wloadvysi
fall
wstorevzsi
w*scale
wloadvzsi
wstorewxsi
w*scale
wloadwxsi
wstorewysi
w*scale
wloadwysi
wstorewzsi
w*scale
wloadwzsi
loopforifixnumfromwide-endbelowcountdo
let
scale
/
+1f0
*h
arefdampingsi
declare
single-floatscale
setf
arefvxsi
*scale
arefvxsi
arefvysi
+
*scale
arefvysi
if
plusp
arefinverse-massesi
gravity0f0
arefvzsi
*scale
arefvzsi
arefwxsi
*scale
arefwxsi
arefwysi
*scale
arefwysi
arefwzsi
*scale
arefwzsi
awake
defmethodphysics-integrate-positions
kernels
eql,family
awakeh

The delta lanes go four at a time; the orientations, whose

normalization the eye alone consumes, stay scalar.

declare
single-floath
records:with-columnar-buffer-storage
countrow
vxsvx
vysvy
vzsvz
dxsdx
dysdy
dzsdz
awakephysics-body-columns
declare
ignorerow
optimize
speed3
safety0
let
wide-end
*4
floorcount4
hs
wsplath
declare
fixnumwide-end
loopforifixnumfrom0belowwide-endby4do
wstoredxsi
w+
wloaddxsi
w*hs
wloadvxsi
wstoredysi
w+
wloaddysi
w*hs
wloadvysi
wstoredzsi
w+
wloaddzsi
w*hs
wloadvzsi
loopforifixnumfromwide-endbelowcountdo
setf
arefdxsi
+
arefdxsi
*h
arefvxsi
arefdysi
+
arefdysi
*h
arefvysi
arefdzsi
+
arefdzsi
*h
arefvzsi
awake
defmethodphysics-warm-start
kernels
eql,family
constraintsawakestarts
do-physics-colors
startendstarts
%wide-physics-warm-startconstraintsawakestartend
constraints
defun,
intern
formatnil"%~A-PHYSICS-WARM-START"family
constraintsawakestartend
declare
fixnumstartend
let
wide-end
+start
*4
floor
-endstart
4
declare
fixnumwide-end
with-physics-kernel-columns
constraintsawake
declare
optimize
speed3
safety0
loopforcfixnumfromstartbelowwide-endby4do
let*
c1
+c1
c2
+c2
c3
+c3
ia0
arefbody-asc
ia1
arefbody-asc1
ia2
arefbody-asc2
ia3
arefbody-asc3
ib0
arefbody-bsc
ib1
arefbody-bsc1
ib2
arefbody-bsc2
ib3
arefbody-bsc3
nx
wloadnxsc
ny
wloadnysc
nz
wloadnzsc
lambda-n
wloadnormal-impulsesc
f1
wloadtangent-impulses-1c
f2
wloadtangent-impulses-2c
px
w+
w+
w*lambda-nnx
w*f1
wloadt1xsc
w*f2
wloadt2xsc
py
w+
w+
w*lambda-nny
w*f1
wloadt1ysc
w*f2
wloadt2ysc
pz
w+
w+
w*lambda-nnz
w*f1
wloadt1zsc
w*f2
wloadt2zsc
rax
wloadraxsc
ray
wloadraysc
raz
wloadrazsc
rbx
wloadrbxsc
rby
wloadrbysc
rbz
wloadrbzsc
ima
wgatherinverse-massesia0ia1ia2ia3
imb
wgatherinverse-massesib0ib1ib2ib3
iia
wgatherinverse-inertiasia0ia1ia2ia3
iib
wgatherinverse-inertiasib0ib1ib2ib3
rx
wloadrolling-impulses-xc
ry
wloadrolling-impulses-yc
rz
wloadrolling-impulses-zc
declare
fixnumc1c2c3ia0ia1ia2ia3ib0ib1ib2ib3
wscattervxsia0ia1ia2ia3
w-
wgathervxsia0ia1ia2ia3
w*imapx
wscattervysia0ia1ia2ia3
w-
wgathervysia0ia1ia2ia3
w*imapy
wscattervzsia0ia1ia2ia3
w-
wgathervzsia0ia1ia2ia3
w*imapz
wscattervxsib0ib1ib2ib3
w+
wgathervxsib0ib1ib2ib3
w*imbpx
wscattervysib0ib1ib2ib3
w+
wgathervysib0ib1ib2ib3
w*imbpy
wscattervzsib0ib1ib2ib3
w+
wgathervzsib0ib1ib2ib3
w*imbpz
wscatterwxsia0ia1ia2ia3
w-
wgatherwxsia0ia1ia2ia3
w*iia
w+
w-
w*raypz
w*razpy
rx
wscatterwysia0ia1ia2ia3
w-
wgatherwysia0ia1ia2ia3
w*iia
w+
w-
w*razpx
w*raxpz
ry
wscatterwzsia0ia1ia2ia3
w-
wgatherwzsia0ia1ia2ia3
w*iia
w+
w-
w*raxpy
w*raypx
rz
wscatterwxsib0ib1ib2ib3
w+
wgatherwxsib0ib1ib2ib3
w*iib
w+
w-
w*rbypz
w*rbzpy
rx
wscatterwysib0ib1ib2ib3
w+
wgatherwysib0ib1ib2ib3
w*iib
w+
w-
w*rbzpx
w*rbxpz
ry
wscatterwzsib0ib1ib2ib3
w+
wgatherwzsib0ib1ib2ib3
w*iib
w+
w-
w*rbxpy
w*rbypx
rz
when
<wide-endend
%physics-warm-start-scalarconstraintsawakewide-endend
constraints
defmethodphysics-solve-contacts
kernels
eql,family
constraintsawakestartsinv-huse-bias-pelapsedpush-max
declare
single-floatinv-helapsedpush-max
do-physics-colors
startendstarts
%wide-physics-solve-contactsconstraintsawakestartendinv-huse-bias-pelapsedpush-max
constraints
defun,
intern
formatnil"%~A-PHYSICS-SOLVE-CONTACTS"family
constraintsawakestartendinv-huse-bias-pelapsedpush-max
declare
fixnumstartend
single-floatinv-helapsedpush-max
let
wide-end
+start
*4
floor
-endstart
4
declare
fixnumwide-end
with-physics-kernel-columns
constraintsawake
declare
optimize
speed3
safety0
let
zero
wsplat0f0
one
wsplat1f0
inv-hs
wsplatinv-h
elapseds
wsplatelapsed
push-maxs
wsplat
-push-max
epsilon
wsplat1e-12
macrolet
solve-lane-groups
biasing-p
`
loopforcfixnumfromstartbelowwide-endby4do
let*
c1
+c1
c2
+c2
c3
+c3
ia0
arefbody-asc
ia1
arefbody-asc1
ia2
arefbody-asc2
ia3
arefbody-asc3
ib0
arefbody-bsc
ib1
arefbody-bsc1
ib2
arefbody-bsc2
ib3
arefbody-bsc3
nx
wloadnxsc
ny
wloadnysc
nz
wloadnzsc
rax
wloadraxsc
ray
wloadraysc
raz
wloadrazsc
rbx
wloadrbxsc
rby
wloadrbysc
rbz
wloadrbzsc
kvx
wloadkvxsc
kvy
wloadkvysc
kvz
wloadkvzsc
ima
wgatherinverse-massesia0ia1ia2ia3
imb
wgatherinverse-massesib0ib1ib2ib3
iia
wgatherinverse-inertiasia0ia1ia2ia3
iib
wgatherinverse-inertiasib0ib1ib2ib3
vax
wgathervxsia0ia1ia2ia3
vay
wgathervysia0ia1ia2ia3
vaz
wgathervzsia0ia1ia2ia3
wax
wgatherwxsia0ia1ia2ia3
way
wgatherwysia0ia1ia2ia3
waz
wgatherwzsia0ia1ia2ia3
vbx
wgathervxsib0ib1ib2ib3
vby
wgathervysib0ib1ib2ib3
vbz
wgathervzsib0ib1ib2ib3
wbx
wgatherwxsib0ib1ib2ib3
wby
wgatherwysib0ib1ib2ib3
wbz
wgatherwzsib0ib1ib2ib3
dpx
w+
w-
wgatherdxsib0ib1ib2ib3
wgatherdxsia0ia1ia2ia3
w*kvxelapseds
dpy
w+
w-
wgatherdysib0ib1ib2ib3
wgatherdysia0ia1ia2ia3
w*kvyelapseds
dpz
w+
w-
wgatherdzsib0ib1ib2ib3
wgatherdzsia0ia1ia2ia3
w*kvzelapseds
s
w+
w+
w+
wloadseparationsc
w*dpxnx
w*dpyny
w*dpznz
speculative
w>szero
bias
wblendspeculative
w*sinv-hs
,
ifbiasing-p'
wmax
w*
wloadbias-ratesc
s
push-maxs
'zero
mass-scale
wblendspeculativeone,
ifbiasing-p'
wloadmass-scalesc
'one
impulse-scale
wblendspeculativezero,
ifbiasing-p'
wloadimpulse-scalesc
'zero
declare
fixnumc1c2c3ia0ia1ia2ia3ib0ib1ib2ib3
with-relative-velocities
vraxvrayvrazvrbxvrbyvrbz
let*
vn
w+
w+
w*
w-vrbxvrax
nx
w*
w-vrbyvray
ny
w*
w-vrbzvraz
nz
old
wloadnormal-impulsesc
delta
w-
w-zero
w*
wloadnormal-massesc
w+
w*mass-scalevn
bias
w*impulse-scaleold
new
wmax
w+olddelta
zero
applied
w-newold
px
w*appliednx
py
w*appliedny
pz
w*appliednz
wstorenormal-impulsescnew
wstoretotal-normal-impulsesc
w+
wloadtotal-normal-impulsesc
new
apply-impulsepxpypz
,@
unlessbiasing-p'

Rolling resistance.

let*
resistance
wloadrolling-resistancesc
rolling-mass
wloadrolling-massesc
ox
wloadrolling-impulses-xc
oy
wloadrolling-impulses-yc
oz
wloadrolling-impulses-zc
nx2
w-ox
w*rolling-mass
w-wbxwax
ny2
w-oy
w*rolling-mass
w-wbyway
nz2
w-oz
w*rolling-mass
w-wbzwaz
limit
w*resistancenew
magnitude-squared
w+
w+
w*nx2nx2
w*ny2ny2
w*nz2nz2
over
w>magnitude-squared
w+
w*limitlimit
epsilon

Where nothing is over, the divisor is

pushed to one and the quotient discarded.

scale
wblendover
w/limit
wsqrt
wblendovermagnitude-squaredone
one

A resistance of zero keeps its impulse at zero.

active
w>resistancezero
nx3
wblendactive
w*nx2scale
ox
ny3
wblendactive
w*ny2scale
oy
nz3
wblendactive
w*nz2scale
oz
wstorerolling-impulses-xcnx3
wstorerolling-impulses-ycny3
wstorerolling-impulses-zcnz3
let
ax
w-nx3ox
ay
w-ny3oy
az
w-nz3oz
setfwax
w-wax
w*iiaax
way
w-way
w*iiaay
waz
w-waz
w*iiaaz
wbx
w+wbx
w*iibax
wby
w+wby
w*iibay
wbz
w+wbz
w*iibaz

Friction.

let*
t1x
wloadt1xsc
t1y
wloadt1ysc
t1z
wloadt1zsc
t2x
wloadt2xsc
t2y
wloadt2ysc
t2z
wloadt2zsc
with-relative-velocities
vraxvrayvrazvrbxvrbyvrbz
let*
vrx
w-vrbxvrax
vry
w-vrbyvray
vrz
w-vrbzvraz
vt1
w+
w+
w*vrxt1x
w*vryt1y
w*vrzt1z
vt2
w+
w+
w*vrxt2x
w*vryt2y
w*vrzt2z
tangent-mass
wloadtangent-massesc
o1
wloadtangent-impulses-1c
o2
wloadtangent-impulses-2c
n1
w-o1
w*tangent-massvt1
n2
w-o2
w*tangent-massvt2
limit
w*
wloadfrictionsc
new
length-squared
w+
w*n1n1
w*n2n2
over
w>length-squared
w*limitlimit
scale
wblendover
w/limit
wsqrt
wblendoverlength-squaredone
one
n1
w*n1scale
n2
w*n2scale
wstoretangent-impulses-1cn1
wstoretangent-impulses-2cn2
let*
d1
w-n1o1
d2
w-n2o2
px
w+
w*d1t1x
w*d2t2x
py
w+
w*d1t1y
w*d2t2y
pz
w+
w*d1t1z
w*d2t2z
apply-impulsepxpypz
wscattervxsia0ia1ia2ia3vax
wscattervysia0ia1ia2ia3vay
wscattervzsia0ia1ia2ia3vaz
wscatterwxsia0ia1ia2ia3wax
wscatterwysia0ia1ia2ia3way
wscatterwzsia0ia1ia2ia3waz
wscattervxsib0ib1ib2ib3vbx
wscattervysib0ib1ib2ib3vby
wscattervzsib0ib1ib2ib3vbz
wscatterwxsib0ib1ib2ib3wbx
wscatterwysib0ib1ib2ib3wby
wscatterwzsib0ib1ib2ib3wbz
ifuse-bias-p
solve-lane-groupst
solve-lane-groupsnil
when
<wide-endend
%physics-solve-contacts-scalarconstraintsawakewide-endendinv-huse-bias-pelapsedpush-max
constraints
defmethodphysics-apply-restitution
kernels
eql,family
constraintsawakestartsthreshold
declare
single-floatthreshold
do-physics-colors
startendstarts
%wide-physics-apply-restitutionconstraintsawakestartendthreshold
constraints
defun,
intern
formatnil"%~A-PHYSICS-APPLY-RESTITUTION"family
constraintsawakestartendthreshold
declare
fixnumstartend
single-floatthreshold
let
wide-end
+start
*4
floor
-endstart
4
declare
fixnumwide-end
with-physics-kernel-columns
constraintsawake
declare
optimize
speed3
safety0
let
zero
wsplat0f0
thresholds
wsplat
-threshold
loopforcfixnumfromstartbelowwide-endby4do
let*
c1
+c1
c2
+c2
c3
+c3
ia0
arefbody-asc
ia1
arefbody-asc1
ia2
arefbody-asc2
ia3
arefbody-asc3
ib0
arefbody-bsc
ib1
arefbody-bsc1
ib2
arefbody-bsc2
ib3
arefbody-bsc3
relative
wloadrelative-velocitiesc
restitution
wloadrestitutionsc
total
wloadtotal-normal-impulsesc
nx
wloadnxsc
ny
wloadnysc
nz
wloadnzsc
rax
wloadraxsc
ray
wloadraysc
raz
wloadrazsc
rbx
wloadrbxsc
rby
wloadrbysc
rbz
wloadrbzsc
kvx
wloadkvxsc
kvy
wloadkvysc
kvz
wloadkvzsc
ima
wgatherinverse-massesia0ia1ia2ia3
imb
wgatherinverse-massesib0ib1ib2ib3
iia
wgatherinverse-inertiasia0ia1ia2ia3
iib
wgatherinverse-inertiasib0ib1ib2ib3
vax
wgathervxsia0ia1ia2ia3
vay
wgathervysia0ia1ia2ia3
vaz
wgathervzsia0ia1ia2ia3
wax
wgatherwxsia0ia1ia2ia3
way
wgatherwysia0ia1ia2ia3
waz
wgatherwzsia0ia1ia2ia3
vbx
wgathervxsib0ib1ib2ib3
vby
wgathervysib0ib1ib2ib3
vbz
wgathervzsib0ib1ib2ib3
wbx
wgatherwxsib0ib1ib2ib3
wby
wgatherwysib0ib1ib2ib3
wbz
wgatherwzsib0ib1ib2ib3
declare
fixnumc1c2c3ia0ia1ia2ia3ib0ib1ib2ib3
with-relative-velocities
vraxvrayvrazvrbxvrbyvrbz
let*
vn
w+
w+
w*
w-vrbxvrax
nx
w*
w-vrbyvray
ny
w*
w-vrbzvraz
nz
old
wloadnormal-impulsesc
delta
w-zero
w*
wloadnormal-massesc
w+vn
w*restitutionrelative

Only a real hit bounces: it pushed, and it

arrived fast enough, and it has bounce in it.

bouncing
,
sym"U32.4-AND"
,
sym"U32.4-AND"
w>restitutionzero
w<relativethresholds
w>totalzero
delta
wblendbouncingdeltazero
new
wmax
w+olddelta
zero
applied
w-newold
px
w*appliednx
py
w*appliedny
pz
w*appliednz
wstorenormal-impulsescnew
wstoretotal-normal-impulsesc
wblendbouncing
w+totalnew
total
apply-impulsepxpypz
wscattervxsia0ia1ia2ia3vax
wscattervysia0ia1ia2ia3vay
wscattervzsia0ia1ia2ia3vaz
wscatterwxsia0ia1ia2ia3wax
wscatterwysia0ia1ia2ia3way
wscatterwzsia0ia1ia2ia3waz
wscattervxsib0ib1ib2ib3vbx
wscattervysib0ib1ib2ib3vby
wscattervzsib0ib1ib2ib3vbz
wscatterwxsib0ib1ib2ib3wbx
wscatterwysib0ib1ib2ib3wby
wscatterwzsib0ib1ib2ib3wbz
when
<wide-endend
%physics-apply-restitution-scalarconstraintsawakewide-endendthreshold
constraints

The families this image can run. NEON is the whole 128-bit menu on arm64; SSE is the same width on x86-64. Availability is a runtime property (#C3F7XM): the methods are defined whenever the package exists, and the default family is chosen at load time.

#+arm64 (define-wide-physics-kernels :neon #:sb-simd-neon :blend :bit-select)
#+x86-64
define-wide-physics-kernels:sse2#:sb-simd-sse2:blend:and-or
defunavailable-physics-kernel-families

The kernel families this image can run, fastest first.

append#+arm64(and (physics-instruction-set-available-p :neon) '(:neon))#+x86-64'
:scalar