Crew 3D Downlink

Technical proposal · BAS on-board activity recognition ·

Crew 3D Downlink

Perception and downlink for crew science on Bharatiya Antariksh Station: what we built, what we measured, and where we propose to go next.

Abstract. On a station without continuous ground support, controllers still need to know what the crew is doing, with which hardware, and whether a procedure is on track. We propose to downlink state rather than video: each crew member as body-model parameters, the module as geometry, and a short language log written on board. On public NASA footage from inside the ISS, sending body-model parameters plus a small scene image (about per second) gives the ground the same crew meshes as sending full frames (about per second). We report where learned end-to-end compression and split inference currently fail, and set out distillation, JEPA-style predictive latents and a module digital twin as the route to on-board operation at around 1 KB per second.

  1. 1Problem and requirements
  2. 2System architecture
  3. 3Experiments so far
  4. 4Proposed directions
  5. 5Experiment plan
  6. 6Limitations
  7. References

1Problem and requirements

The BAS challenge asks for a model that recognises and validates the sequence of a predefined experiment on board. Validation depends on physical relations (which hand holds which sample, whether a glovebox port is in use), so we treat 3D crew state as the primary signal and video as a fallback. Table 1 lists what the system must deliver and how each requirement is measured in this report.

Table 1. Requirements and how we measure them.
RequirementWhy it matters on BASMeasured as
Crew 3D pose relative to hardwareStep validation depends on hand–object and body–rack relations.Joint error (MPJPE) and aligned pose error (PA-MPJPE), mm
Scene and hardware geometryControllers need to see where crew and equipment are.Relative depth error; shape error after scale correction
Downlink budgetLinks are shared and not always available; video cannot be assumed.Bytes per second at one frame per second
On-board computeFlight hardware trails ground GPUs by years.Parameters; ms per frame on 4 CPU cores
Orientation invarianceIn microgravity there is no floor; crew work at any angle to the rack.Planned: error versus body orientation (E9)
Trustworthy reportingAlerts must come from measured state, not a language model's guess.Planned: step accuracy and error-detection rate (E10)

2System architecture

Figure 1 shows the pipeline. On board, perception turns each frame into a state packet. The downlink is tiered: a text log always, the state packet at one frame per second, a scene image when the view changes, and full frames only on request. The ground rebuilds the 3D scene from whatever arrives.

ON BOARD GROUND downlink Camera, 1 fps Person detectorboxes SAM 3D Bodystudent later (N3) Scene encodersmall JPEG today Vision-language logQwen3-VL-2B crops State packet MHR params / person camera FOV step state (planned) scene keyframe T0 text logalways · ~20 B/s T1 state packetevery second · <1 KB/person T2 scene keyframeon view change · ~15 KB T3 full frameon request · ~160 KB MHR body modelexact meshes Scene rebuildMoGe-2 or twin (N2) 3D view +ops console bodies placed at scene depth
Figure 1. Pipeline and downlink tiers. Teal is the tier that carries the crew: body-model parameters every second. Solid arrows exist in the current prototype; the dashed path (full frames on request) and the items marked N2 and N3 are proposed in Section 4.

3Experiments so far

3.1 Data and reference models

We use ten public NASA research videos from inside the ISS ( frames at one or two frames per second). The last part of each video is held out. Two large models run on the full-quality frames and serve as references: MoGe-2 for metric scene geometry, and SAM 3D Body for each crew member as parameters of the MHR body model. No model was fine-tuned for microgravity. All numbers below are measured on held-out frames.

3.2 E1: Full frames against body parameters plus a small image

The simplest useful change is to stop sending pixels for the crew. In the parameter mode, SAM 3D Body runs on board and each crew member goes down as its MHR parameters in float16: per person per frame, plus once per person for body shape. The module goes down as a 320-pixel JPEG. The ground rebuilds meshes from the parameters exactly and rebuilds the module with MoGe-2. Figure 2 compares both modes on a 51-second held-out segment.

0 s

Camera frame (on board)

Captured at one frame per second; never sent in parameter mode.

Camera frame at the current time

Mission log (T0)

    Ground 3D from full frames

    Ground 3D from parameters + small image

    Figure 2. Interactive: 51 s of held-out NASA footage (glovebox and sample handling) at one frame per second. Left 3D view: the ground receives each frame as a JPEG and runs MoGe-2 and SAM 3D Body. Right: the ground receives body-model parameters and a 320-pixel JPEG. Drag either view to rotate both. Grey meshes: full frames; teal: parameters.
    Finding. Parameter mode needs fewer bytes than full frames for the same crew meshes. About 95% of its budget is the scene image, which points to the next saving: send the scene only when it changes (E6).

    3.3 E2: Reconciling body and scene scale

    SAM 3D Body and MoGe-2 each estimate distance from a single image, and they disagree. For every crew member we compared the mesh's visible surface with MoGe-2's depth at the same pixels, then scaled the mesh about the camera centre until they matched. Scaling about the camera keeps the mesh's outline in the image unchanged. Figure 3 shows the correction each person needed.

    Figure 3. Scale correction applied to each crew mesh so its surface meets the scene depth.

    3.4 E3: Split inference at SAM 3D Body's image embedding

    Could the station run only SAM 3D Body's image backbone and send its internal embedding, leaving the decoder on the ground? We patched the model at that point, compressed the embedding (float16, 8-bit, or PCA to k channels plus 8-bit, with the basis fitted once on training frames), and compared the ground's output with uncompressed SAM 3D Body on the same held-out frames and boxes.

    Table 2. Split inference: bytes per person against error relative to uncompressed SAM 3D Body. Each person needs three crops (body and both hands).
    What is sentPer personVariance keptJoint errorPose erroralignedPosition error
    Negative result. The embedding is only accurate at sizes larger than a full-frame JPEG. Its information is spread across many channels (256 principal components keep 90% of the variance), and the decoder is sensitive to what is lost: position error grows from centimetres to over a metre. SAM 3D Body's most compact representation is its own output, the parameter vector, which is what E1 sends.

    3.5 E4: A learned end-to-end latent

    We trained our own encoder and decoder to reproduce both reference models from one latent. The encoder is a frozen DINOv2 ViT-S/14 with a small trained bottleneck (); the latent travels with a 32 × 18 colour thumbnail. The ground decoder predicts scene depth and a heatmap of crew centres with MHR parameters at each peak (CenterNet-style). Table 3 compares it with sending a small JPEG and running MoGe-2 on the ground.

    Table 3. Learned latent against JPEG plus ground-side MoGe-2, on the same held-out frames. Scene errors are relative depth error before and after correcting overall distance. Crew errors are for crew members the decoder found.
    DownlinkBytesper frameDepth errorShape errorColourPSNRCrew foundJoint errorPose erroraligned

    Scene

    Crew

    Learned latentJPEG + MoGe-2 on the ground
    Figure 4. Error against bytes per frame (log scale).
    Negative result, with a clear cause. At about 2 KB the latent beats a JPEG of the same size on absolute distance, and it carries the crew, but its crew placement is off by roughly 20 cm per joint. Quality barely changes between 2 KB and 11 KB: the model is limited by training data (about 1,300 edited frames), not by bytes. That motivates distillation on more, and synthetic, data (N3) rather than a larger latent.

    3.6 E5: An on-board mission log

    A 2-billion-parameter vision-language model (Qwen3-VL-2B) reads each 10-second window and writes a short entry for controllers, sent as text (tier T0). Two design rules made the log usable. First, counting is delegated: the model is given the person detector's crew counts instead of counting across ten images itself. Second, it is told what it must not assume: crew float, and empty spacesuits are equipment. Its entries are visible in Figure 2. The log describes; it does not validate. Step validation belongs to a checker over measured state (N7).

    4Proposed directions

    Each direction below states the idea, why it matters for BAS, and the experiment that would confirm or reject it. Section 5 orders them into a plan.

    N1Event-triggered state downlink

    Send the crew state every second, but the scene only when it changes. The ground and station run the same simple predictor (the last keyframe); the station sends a new keyframe when its view differs from what the ground would predict. Bandwidth then scales with activity rather than time.

    For BAS
    About 1–2 KB/s for continuous crew monitoring, low enough for a shared or intermittent link.
    Test
    E6: bytes per second and scene error against a keyframe threshold.

    N2Module digital twin on the ground

    The station's interior is designed, built and documented. Its geometry can live on the ground as a 3D model, registered to each camera once. The downlink then carries only what moves: crew, and hardware whose state changes.

    For BAS
    A complete 3D scene, not a single-view shell, at almost no downlink cost.
    Test
    Register a CAD stand-in of a rack and glovebox to the camera; compare with MoGe-2 geometry.

    N3Distilled on-board perception

    SAM 3D Body has 840M parameters, too many for flight hardware. We distil it into a small student that predicts the same MHR parameters, trained with losses on parameters, on the rebuilt mesh, and on intermediate features. Training data comes from the teacher on real footage plus our synthetic generator, which renders crew at every orientation with exact labels.

    For BAS
    Parameter-mode downlink with a model that fits on board.
    Test
    E8: student error against SAM 3D Body; CPU latency.

    N4Predictive latent state (V-JEPA 2)

    V-JEPA 2 learns video representations by predicting the embeddings of future frames rather than their pixels. On board, the gap between what it predicts and what it then sees is a direct signal of the unexpected: a skipped step, a dropped object, an unusual posture. The same predictor, shared with the ground, turns the downlink into predictive coding: send only the residual.

    For BAS
    Anomaly and deviation detection without labelled failure data.
    Test
    E11: prediction error on correct against deliberately wrong protocol runs.

    N5Language-grounded step recognition (VL-JEPA)

    VL-JEPA predicts the embedding of a text answer instead of generating tokens, and decodes to words only when needed. Protocol steps written as sentences ("open the glovebox port", "transfer sample A") become targets in the same embedding space. A new experiment then only needs its protocol text, not new training, and the log is decoded only when the recognised step changes.

    For BAS
    Open-vocabulary step recognition, and less on-board decoding work.
    Test
    E10: step accuracy on the synthetic protocol, zero-shot against trained.

    N6Rack-frame, orientation-free body state

    In microgravity "upright" is meaningless. We express each crew member's pose in the rack's coordinate frame, not the camera's or gravity's, and train and test with bodies at every orientation from the synthetic generator.

    For BAS
    Step checks that hold for inverted and sideways crew.
    Test
    E9: reference-model error against body orientation; with and without canonicalisation.

    N7Validation from measured state

    A protocol checker compares the sequence of measured states (hand contacts, object positions, lid open or closed) against the expected procedure and raises explicit alerts. The language model narrates the checker's output; it never decides whether a step was valid.

    For BAS
    Alerts that can be traced to a measurement.
    Test
    E10: error-detection rate and latency on runs with injected mistakes.

    N8Task-oriented feature codecs

    E3 shows that generic compression of a large model's features fails. Codecs trained for the task, with a learned bottleneck placed where the task loss is measured, compress far better in published edge–cloud work. We would train one at the student's feature layer, with the ground-side decoder fine-tuned jointly.

    For BAS
    A fallback when the student's own output is not enough.
    Test
    Bytes against joint error at the student's features, compared with E3.

    5Experiment plan

    Table 4. Experiments, with the result that would count as success.
    IDQuestionSuccess criterionStatus
    E1Do body parameters plus a small image give the ground the same crew as full frames?Identical meshes at ≥ 5× fewer bytesdone
    E2Do body and scene models agree on scale, and can they be reconciled?Mesh surface meets scene depth; plausible body heightsdone
    E3Is SAM 3D Body's internal embedding a good thing to send?Accurate below full-frame sizedone: no
    E4Can one learned latent carry scene and crew?Crew within 10 cm per joint at ≤ 5 KBdone: not yet
    E5Can a small on-board model write a useful log?Correct crew counts and hardware in each entrydone
    E6How low can event-triggered keyframes take the downlink? (N1)≤ 2 KB/s with the same crew meshesnext
    E7How small can the scene image be in parameter mode?Shape error within 2 points of the 320-pixel imagenext
    E8Can a distilled student replace SAM 3D Body on board? (N3)Aligned pose error ≤ 60 mm; ≤ 150 ms on 4 CPU coresnext
    E9How much does body orientation hurt the reference model? (N6)Error curve by orientation; fix recovers upright accuracyplanned
    E10Step recognition and validation on the synthetic protocol (N5, N7)≥ 95% step accuracy; every injected mistake flaggedplanned
    E11Does predictive error flag protocol deviations? (N4)Separates correct from wrong runs without labelsplanned

    6Limitations

    References

    1. Meta AI. SAM 3D Body: robust full-body human mesh recovery. arXiv 2602.15989; code.
    2. Meta AI. MHR: Momentum Human Rig. arXiv 2511.15586.
    3. Wang et al. MoGe-2: accurate monocular geometry with metric scale. code.
    4. Oquab et al. DINOv2: learning robust visual features without supervision. arXiv 2304.07193.
    5. Qwen Team. Qwen3-VL. model card.
    6. Meta AI. V-JEPA 2: self-supervised video models enable understanding, prediction and planning (2025). paper.
    7. Chen et al. VL-JEPA: joint embedding predictive architecture for vision-language. arXiv 2512.10942.
    8. Task-oriented communication for human action understanding via edge-cloud co-inference. arXiv 2605.07354.
    9. Zhou, Wang, Krähenbühl. Objects as points (CenterNet). arXiv 1904.07850.
    10. NASA Image and Video Library: ISS crew research videos (public domain). images.nasa.gov.