Skip to content

Solve the surface, not the skeleton ​

By Oleg Sidorkin, CTO and Co-Founder of Cinevva

Part 1 measured four auto-riggers against artist rigs on twelve CC0 animals and found one shared flaw: every one of them builds a quadruped with nearly straight legs. It ended with a plan to repair rigs: with bone lengths fixed, bending a straight leg pulls the foot toward the hip.

For the next experiment, we used the animated surface as the target for each rig.

A skeleton is a machine for deforming a surface. The surface is the thing the player sees, and the artist clip already tells us exactly where that surface should be, vertex by vertex, on every frame. So skip the skeleton-to-skeleton mapping entirely. Play the artist clip once, record the deformed mesh, then solve the auto-rigged skeleton's transforms so that its own skinning reproduces that surface as closely as it can. The result is a clip that lives natively in the new rig's space. This also works with anonymous bone names such as bone_0.

Pipeline diagram: artist clip to deformed surface to per-frame solve to a clip in the rig's own space
The pipeline matches deformed surfaces and solves bone transforms. It requires correspondence between the meshes, but no bone-to-bone mapping.

This uses the transform-fitting part of skinning decomposition, as in SSDR from Le and Deng in 2012. Motion capture solvers also fit rigs to dense markers. Here, the rigger already supplied the skeleton and the weights, so the only unknowns are the per-frame bone transforms.

The method and control run ​

Per frame, we look for one transform per bone that minimizes the distance between the rig's skinned surface and the target surface. With weights fixed, each bone's best transform has a closed-form answer, a weighted rigid fit by Horn's quaternion method. Linear blend skinning couples bones through shared vertices, so we sweep the bones in turn, each solve accounting for what every other bone currently contributes, until the residual stops moving. Each frame warm-starts from the previous one.

Everything runs in bind space, where the skinning identity holds exactly whatever conventions the file was exported with. Correspondence between the artist mesh and the rigger's re-welded or subdivided mesh is nearest-neighbour on the aligned rest poses, which is exact here because both are the same animal.

We first solve the artist rig against its own clip, where a perfect answer exists by construction. Residual error in this control helps identify problems in the solver or target preparation. The control lands at 0.008% of the model diagonal, with the single worst vertex at 0.08%. We use that as the baseline for the other rigs.

Three problems in the control run ​

The first control run plateaued at 3%, even with thirty times more iterations. We found three causes.

The largest error came from our use of the Three.js API. After the change from boneTransform to applyBoneTransform, our call needed to supply the vertex position in the input vector. Fed a reused scratch vector, it computes each vertex from the previous vertex's output. The resulting targets looked smooth, which initially led the investigation toward armature scale. Switching to getVertexPosition corrected the target positions.

The other two were smaller. Power iteration on Horn's matrix stalls at about a thousandth of a degree because the spectral shift that makes it converge also crushes the gap it converges along, so we switched to Jacobi rotations, which are exact for a 4x4. And an early-exit test based on relative progress froze the sweeps while they were still grinding downward, so the exit now measures absolute progress against the model's scale.

What each rigger can actually carry ​

With the solver validated, we solved two artist clips, Walk and Gallop, into every rig for the six animals all riggers cover. We call the score transfer quality: the share of the motion's surface displacement the rig reproduces.

Dot plot of transfer quality by rigger for Walk and Gallop, artist control near 100 percent, RigNet near 50
Quality = 1 − solved residual / motion magnitude, averaged over six animals. The control row is the artist rig solved against its own clip.
rigWalkGallop
artist (control)99.8%99.9%
Anything World94.4%91.9%
Tripo92.3%89.1%
SkinTokens, given the artist skeleton82.5%89.4%
SkinTokens80.9%88.0%
RigNet48.5%58.3%

The ranking differs from the fold-ratio comparison in part 1.

RigNet had the best leg fold of the four and was the only rigger whose skeletons passed our drivability gate on all twelve animals. Its Walk transfer-quality score here is 48.5%. Its worst vertices sit 20 to 31% of the model away from where the surface should be, and the failure is consistent across all six animals, 38 to 59% quality on every one. Meanwhile Tripo and Anything World, the two straightest-legged rigs we measured, carry animation best.

The solved clips include joint translations as well as rotations. With translation available, a straight chain can still track a folding surface, so the fold stops being the binding constraint. What binds instead is whether the skinning weights carve the mesh into pieces that follow bones cleanly. In these runs, Tripo and Anything World reproduced the surface more closely. RigNet retained large residuals as the solver adjusted its irregular chains.

So the two measurements answer two different questions. Fold ratio governs what a rotation-driven runtime rig can pose, which is part 1's world. Transfer quality governs what a baked, solved clip can express, which is this one's. Tripo, the rigger we ship, reproduced 92.3% of the Walk surface displacement in this test.

Supplying the artist skeleton to SkinTokens moved Walk transfer quality from 80.9% to 82.5% and Gallop from 88.0% to 89.4%. The small improvement suggests that skinning weights also limit the result.

Inspecting the surface error ​

The same solver drives a live viewer where every rig's full mesh is skinned by its solved transforms, drawn over the artist's ground-truth surface as a dark silhouette, and every vertex is coloured by its momentary error. Blue is zero. Red is 3% of the model diagonal or worse.

Walk, solved into every rig. Left to right: artist ground truth, Tripo, Anything World, SkinTokens, RigNet. RigNet's legs show large red regions and deviate visibly from the reference silhouette.
Gallop comparison. The hind legs at full extension show the differences in surface alignment, with larger deviations on RigNet.

The viewer uses the same skinning function as the measurements. The SkinTokens deer has 181,000 vertices, above our 60,000-vertex limit for full-mesh correspondence, so its panel shows a silhouette comparison without a heatmap.

What this changes for us ​

A solve takes half a second to a few seconds per clip per rig, in browser JavaScript, no GPU involved. That's cheap enough to bake a quadruped's whole clip set when the rig is created. Baked clips make the runtime simpler as well: no skeleton retargeting for animals at all, just clip playback in the rig's own space, with our existing foot locking on top for ground contact.

It also gives us a second acceptance gate. Part 1's geometric checks say whether a rig is drivable. Transfer quality says how much of a real animation it will keep. A rigger can pass the first and fail the second, and RigNet does exactly that.

Limits of the benchmark ​

These scores describe the fitted clips under the solver settings used here. The solver may translate joints, so bone lengths aren't preserved during motion, and a rig leaning hard on translation could read rubbery at extremes. The gallop's worst vertices, 9 to 20% momentarily even on the good rigs, live in exactly those moments. A rotations-only comparison remains to be measured. It would show how much of the straight-leg limitation from part 1 persists with fixed bone lengths.

The production gap is correspondence. In this benchmark the rigger rigged the same mesh the artist animated, so vertex matching is trivial. A user's uploaded animal is a different mesh, and getting artist surface motion onto it first is its own problem, with known routes we haven't measured yet.

And the sample: six animals, two clips, 24 frames per clip, constraints subsampled to five thousand vertices on the densest mesh. The rankings apply to this sample and these solver settings.

The animals are CC0 from Quaternius. To reproduce the comparison, start by solving an artist rig against its own clip and checking the residual. If you've solved animation into rigs before and hit the traps we did, or different ones, we'd like to compare notes in our Discord.