Skip to content

Four auto-riggers, twelve animals, one shared flaw: straight legs ​

By Oleg Sidorkin, CTO and Co-Founder of Cinevva

Take an artist-made walk cycle from a deer, retarget it onto the same deer after an auto-rigger has rigged it, and the result looks wrong in a way that's hard to name. The feet land in roughly the right places. The timing is right. It still reads like a table walking.

The investigation started with our retargeting code. Examining the bind pose revealed nearly straight leg chains, so we measured how much fold each rig had before animation.

The number ​

Take one leg. Add up the lengths of its bone segments, hip to foot. Call that the reach. Now measure the straight-line distance from the hip joint to the foot joint in the bind pose. Call that the span. Divide reach by span.

A real animal standing still has a folded leg, so the chain is meaningfully longer than the line it spans. Across the twelve artist-rigged animals in the Quaternius Ultimate Animated Animals pack, that ratio runs from 1.15 to 1.50, averaging 1.30. A deer's hind leg is 1.40. A fox's front leg is 1.29.

With bone lengths fixed, a leg with no fold can't bend without shortening its span. If the joints sit in a straight line, the only way to flex the knee is to pull the foot toward the hip. So either the foot leaves the ground or the body sinks. Ask that rig for the artist's knee angle and you get a crouch. Rotation-only retargeting cannot preserve both the hip-to-foot span and the bend when the chain has too little length.

Two leg chains between the same hip and foot, one folded measuring 1.31 and one nearly straight measuring 1.04
Both chains start at the same hip and end at the same foot. The folded one is 31% longer than the line it spans, so it has slack to bend into. The straight one has none, and the numbers shown are the measured artist mean and Tripo mean.
Animation of a folded chain and a straight chain moving through the same flex, with span readouts

We measured whether the predicted leg chains had enough fold to reproduce the reference stance.

The test setup ​

Twelve animals, all CC0 from the Quaternius pack, each shipping an artist rig of 38 to 51 joints.

The twelve Quaternius CC0 animals rendered with their own materials, from alpaca to wolf
The subjects, with their own materials and their artist bone counts. We removed the supplied rigs before submitting the meshes and kept those rigs as references.

We submitted the unrigged meshes to each rigger and compared the results with the artist's skeletons.

Mesh connectivity required preparation. Those glTF files split the mesh into five primitives by material, and because the models are hard-edged low-poly, almost every vertex is duplicated per adjacent face. The fox stores 3,752 vertices and has 926 distinct ones. Riggers that build a graph over mesh topology read those coincident-but-unshared seams as cracks, so the head becomes a separate object from the body, the geodesic distance between them goes to infinity, and a joint that should sit in the neck has no reason to go there. Welding by position reconnects those seams before the graph is built. Nine of the twelve came out as a single watertight shell. The bull and cow keep their horns as separate pieces and the stag its antlers, which we left in, because real uploads look like that.

For scoring we used two things. The first is our own runtime function that reduces an arbitrary quadruped skeleton to a canonical structure. It has to find four feet in four distinct quadrants, a spine, and a neck. Rigs that fail this check cannot be driven through our current quadruped runtime path.

That function uses bind geometry because the predicted bone names were inconsistent. Tripo numbers its limb chains by index with no fixed meaning, so chain 0 is the hind pair on its deer and the front pair on its horse. Its left and right labels disagree between rigs: Left sits at x=-0.072 on the deer and x=+0.076 on the horse, both facing the same way. On its fox it labels the spine as a limb chain. These inconsistencies led us to identify the chains from their positions and connections.

The second thing is the fold ratio above. Both measures are scale invariant, which matters because these riggers work at wildly different scales. Tripo normalises to roughly a unit cube, so its fox spans 0.13 by 0.34 by 0.65 against the artist's 0.77 by 2.47 by 6.59.

What came back ​

riggerrigsdrivablemean jointsmean fold
artist (reference)121146.81.299
Anything World6637.81.363 (see below)
RigNet121254.11.146
SkinTokens121132.01.078
Tripo121025.81.045
Six animals across five riggers, with the artist rig's IK control bones drawn in grey and its deforming bones in blue
Six animals, five riggers. Each cell carries the live drivability verdict and the fold range. In the artist column, blue bones deform the mesh and grey ones are IK controls, so that column looks busier than the predicted rigs beside it.
Dot plot of mean fold ratio by rigger against the shaded artist range of 1.15 to 1.50, all four predicted riggers short of it
The shaded band is where the artist rigs sit. Every predicted rigger lands left of it, short of the fold a leg needs, except Anything World, which overshoots for the wrong reason.
The same knee rotation applied to the artist, Tripo, Anything World, SkinTokens and RigNet deer hind legs

Anything World's 1.363 average includes abnormally folded front legs, described below. Excluding that anomaly, its legs measure 1.02 to 1.18.

Nearly straight leg chains appeared across all four systems, including Tripo, which we ship as our Pro tier. The chain structures and other defects differed, so we examined those separately.

Training data dominated by bipeds and neutral rest poses could contribute to this pattern. This benchmark measures the output rigs and doesn't establish the cause of the joint-placement errors.

The artist Shiba Inu also failed our drivability test because the detector selected IK controls as feet. These rigs ship an IK control layer: four pole targets parented to the root and sitting half a body-length in front of the animal, plus four IK chains with no parent at all. Those satisfy a geometric definition of "foot" as well as the real feet do, and on the Shiba Inu they win, giving eight false candidates out of twelve.

We removed the control bones and repeated the analysis to check their effect on the fold comparison. The result changed on exactly one of the twelve artist rigs. The Shiba Inu goes from a broken 2 limbs to a clean 4, with a fold of 1.28 to 1.43 that lands right beside the husky's. The other eleven come out identical, fold ratios included. Excluding the controls corrected the Shiba Inu foot detection while leaving the other fold measurements unchanged. The artist band of 1.15 to 1.50 was never measured on control bones, and the predicted rigs carry no control layer at all: Tripo's skeletons have none, and Anything World's five non-deforming joints are just leaf tip markers.

Predicted rigs being pure FK also cuts the other way. In this one narrow sense an auto-rigged animal is easier for a name-blind runtime to read than the artist's.

The four riggers have quite different failure modes ​

Tripo is the straightest at 1.045, and it uses the fewest joints by a wide margin, 25.8 against the artist's 46.8. But it gets the chain shape right where the others don't. Four of its twelve rigs reproduce the artist's exact 3/3/4/4 split, three joints per front leg and four per hind. For those matching chains, repositioning the existing joints is a possible repair. Two of its rigs fail drivability outright: the fox comes back with two limbs, and so does the stag.

Deer skeleton, artist rig beside the Tripo prediction, both using a 3/3/4/4 leg chain
Tripo gets the structure right and the placement wrong. Its deer uses the artist's exact 3/3/4/4 chain split with half the joints, and still measures 1.00 to 1.05 against the artist's 1.18 to 1.40.

RigNet (SIGGRAPH 2020) has a mean fold of 1.146 and was the only one to produce a drivable rig for all twelve. Its chains are the problem. The artist uses 3/3/4/4 on all twelve animals, front and back consistent. RigNet reproduces that shape exactly once. Three of its rigs have a leg of two joints or fewer, which cannot articulate at all, and four have a leg of seven or more. Its horse is 2/2/8/8: two joints in the front legs, eight in the back. The published critique of RigNet is that it "outputs irregular skeletons", and that's what irregular looks like in practice. A bandwidth sweep won't fix it either, because that parameter controls how many joints there are rather than how they're distributed between limbs.

Horse skeleton, artist rig with 3/3/4/4 chains beside the RigNet prediction with 2/2/8/8 chains
RigNet's horse is 2/2/8/8: two joints in each front leg, eight in each hind. A two-joint leg cannot articulate at all. The artist uses 3/3/4/4 on all twelve animals.

Anything World failed to rig two of the eight animals we submitted: the husky came back with FinetuningError: Incorrect segmentation exception and the cow with a bare Unknown error. The six that did rig are all drivable. All six use an identical 5/5/5/5 chain, the most consistent structure any predicted rigger produced.

The front-leg asymmetry accounts for its elevated mean fold. Hind legs are mirrored exactly on all six: 1.10 and 1.10, 1.11 and 1.11, 1.13 and 1.13, and so on to five decimal places. Front legs break on three of the six, with one leg folded far past anything anatomical while its twin stands nearly straight. The alpaca is 3.60 on the front left against 1.87 on the front right. The horse is 2.47 against 1.43. The deer is 1.43 against 2.13, and note that the extreme side is the right one there, so this is not a consistent left bias.

Alpaca skeleton, artist rig beside the Anything World prediction whose fold range runs to 3.60
The alpaca is the clearest case: Anything World returns a uniform 5/5/5/5 chain, and a fold range of 1.10 to 3.60 across its four legs. The hind pair agrees exactly. The front pair does not.

We re-ran the deer and the horse to rule out our own submission. symmetry is a required upload field, and the stored response never echoes back what was sent. Submitted explicitly symmetric, both reproduced the anomaly. So it survives the parameter that exists to prevent it, on a mesh whose own artist rig is symmetric to three decimals.

SkinTokens (2025, from the VAST-AI group behind Tripo) is the tidiest by a distance. Eleven of twelve have mirrored left and right chain lengths, and not one has a leg of two joints or fewer. It's a well-behaved skeleton that stands too straight, at 1.078, rather than an erratic one.

Two things about SkinTokens worth writing down ​

Our 926-vertex fox crashed the Blender-side server mid-request with no Python traceback. The same fox subdivided to 88,706 vertices rigs in 70 seconds. Their own example model works because it was already dense. Adding a subdivision step to roughly 40k vertices before the call turned a total failure into twelve rigs out of twelve. Subdivision was necessary for the low-poly assets in this test.

The more useful discovery is that it has a mode we'd never run. Passing --use_skeleton conditions it on a skeleton you supply instead of inventing one. It then keeps what you gave it. Every output joint sits on an input joint, mean offset 0.31% of the bounding box diagonal, worst case 0.50%. It renames them all to bone_0 and up on export and culls the ones carrying no weight, so recovering the original names requires an additional step.

That relabelling is recoverable. Match the output joints back to the skeleton you supplied by position and every name comes back. That worked on 100% of joints, on all four animals we tried. And because the skeleton is now yours, so is the stance. Our horse measures 1.01 to 1.02 when SkinTokens picks the skeleton, and 1.11 to 1.36 when it's handed the artist's. The artist's own is 1.18 to 1.37.

Two preconditions before that works. It rejects a multi-root skeleton outright, and these artist rigs have five roots, the body plus four IK chains. The first attempt deleted those extra roots and their chains. Dropping those eight bones leaves 90 to 128 vertices with no influence at all, on every animal in the pack, because FF and FFB are what the artist skinned the feet to. Those bones carry deformation weights. Re-parenting each stray root to its nearest body joint keeps every weight and still gives one root.

Recovering bone names ​

Predicted skeletons often use names such as joint_0 or bone_0. We tested two ways to map them to the roles expected by our runtime.

Names can be recovered, as above, when you supplied the skeleton. Or they can be derived. A geometric analysis already knows which four chains are legs, which quadrant each sits in, and which run is spine and which is neck. Those detected roles provide the labels. We ran it on all twelve RigNet rigs, writing names into a convention our runtime already reads, then analysed them again. Twelve of twelve came back identical through the named path and the geometric one. Same limb count, same quadrants, same chain lengths, same spine, same neck, same fold ratios.

A second number: is the skeleton even symmetric? ​

The Anything World front legs prompted a separate symmetry measurement.

It works without reading a single bone name, which matters when half the output is called bone_0. Reflect every joint across the model's symmetry plane, then measure how far each reflection lands from the nearest real joint, as a percentage of the bounding box diagonal. A perfectly mirrored skeleton scores zero. The mesh bounding box defines the reference plane, keeping it independent of the predicted joint positions.

riggermean mirror error
artist0.01%
RigNet0.08%
SkinTokens0.37%
Anything World1.08%
Tripo1.13%
Dot plot of left/right mirror error by rigger, artist 0.01% through Tripo 1.13%
Measured without reading a bone name, so it applies equally to a skeleton whose joints are called bone_0. The input meshes are symmetric, so each number measures asymmetry introduced in the predicted skeleton.
Five horse skeletons rotating side by side, artist beside four predicted rigs

The artist rigs are mirrored to within rounding, which is the expected result and a useful check that the measure works. RigNet also has low mirror error. It explicitly reflects candidate joints across x=0 before clustering them.

Tripo comes last, though only just, and its spread is what stands out rather than its average. Eight of its twelve rigs land under 0.8%, and then the fox is 15.02% and the stag 14.45%. Those two also failed drivability. Their large asymmetries warrant inspection alongside the inconsistent left/right labels described earlier.

These are small percentages. A fraction of a percent is invisible standing still. It shows up when a walk cycle asks the left and right legs to do the same thing half a period apart.

The artist deer walking with its skeleton drawn over the ghosted body

What we're doing about it ​

For rotation-driven animation, two possible repairs operate on the rig and can be tested without retraining.

Supply the skeleton and let the model skin it. SkinTokens already supports this and preserves what it's given, and for humanoids our production rigger already emits a properly named Mixamo skeleton that could feed straight into it. Or reposition the joints in the rig you already have, which for Tripo means moving intermediate leg joints on a chain whose joint count is already correct.

Changing riggers alone would leave much of the fold deficit. The means for Tripo, SkinTokens and RigNet range from about 1.04 to 1.15, compared with the artist mean of 1.30.

Update: there turned out to be a third route, and it changes the picture enough to earn its own article. Instead of retargeting between skeletons, we now solve the animation into each rig's own space from the deformed surface, and the quality ranking this produces reverses the one above. That story is part 2: solve the surface, not the skeleton.

Benchmark scope ​

Sample sizes are uneven and that limits what the averages mean. RigNet, SkinTokens and Tripo all ran on the full twelve. Anything World ran on eight, of which six produced a rig, and the two that failed did so with errors that name themselves rather than with a bad rig we could measure. Tripo's five original rigs were leftovers from earlier retargeting work, so the other seven were rigged for this benchmark through the same endpoint on the same inputs. Adding them moved its mean fold by 0.004 and its mean mirror error from 1.54% down to 1.13%, showing how the averages changed with the larger sample.

The first two Anything World rigs both showed the front-leg anomaly. The next model came back clean. Across the six successful outputs, three had the anomaly, the hind legs were mirrored, and the more folded front leg could be on either side. The initial pair overstated how consistently the problem occurred.

RigNet's checkpoints came from a third-party mirror rather than the paper's own download, which needs interactive confirmation and rate-limits. The results therefore describe our setup with that checkpoint source. Getting it running at all took seven separate undocumented fixes. One is an interactive timezone prompt that hangs the container build with no error at all. Another is a debug viewer that opens a window and then blocks forever. You can't avoid that one by going headless either, because its voxeliser rasterises through OpenGL, so the process needs a display.

Finally, the fold ratio assumes a neutral standing bind pose. It's the right measure for these twelve animals and it would be the wrong measure for a rig delivered mid-stride.

The animals are CC0 from Quaternius, so anyone can repeat this. If you rig quadrupeds and you've measured something different, we'd like to hear it in our Discord.