Perception and downlink for crew science on Bharatiya Antariksh Station: what we built, what we measured, and where we propose to go next.
Abstract. On a station without continuous ground support, controllers still need to know what the crew is doing, with which hardware, and whether a procedure is on track. We propose to downlink state rather than video: each crew member as body-model parameters, the module as geometry, and a short language log written on board. On public NASA footage from inside the ISS, sending body-model parameters plus a small scene image (about per second) gives the ground the same crew meshes as sending full frames (about per second). We report where learned end-to-end compression and split inference currently fail, and set out distillation, JEPA-style predictive latents and a module digital twin as the route to on-board operation at around 1 KB per second.
The BAS challenge asks for a model that recognises and validates the sequence of a predefined experiment on board. Validation depends on physical relations (which hand holds which sample, whether a glovebox port is in use), so we treat 3D crew state as the primary signal and video as a fallback. Table 1 lists what the system must deliver and how each requirement is measured in this report.
| Requirement | Why it matters on BAS | Measured as |
|---|---|---|
| Crew 3D pose relative to hardware | Step validation depends on hand–object and body–rack relations. | Joint error (MPJPE) and aligned pose error (PA-MPJPE), mm |
| Scene and hardware geometry | Controllers need to see where crew and equipment are. | Relative depth error; shape error after scale correction |
| Downlink budget | Links are shared and not always available; video cannot be assumed. | Bytes per second at one frame per second |
| On-board compute | Flight hardware trails ground GPUs by years. | Parameters; ms per frame on 4 CPU cores |
| Orientation invariance | In microgravity there is no floor; crew work at any angle to the rack. | Planned: error versus body orientation (E9) |
| Trustworthy reporting | Alerts must come from measured state, not a language model's guess. | Planned: step accuracy and error-detection rate (E10) |
Figure 1 shows the pipeline. On board, perception turns each frame into a state packet. The downlink is tiered: a text log always, the state packet at one frame per second, a scene image when the view changes, and full frames only on request. The ground rebuilds the 3D scene from whatever arrives.
We use ten public NASA research videos from inside the ISS ( frames at one or two frames per second). The last part of each video is held out. Two large models run on the full-quality frames and serve as references: MoGe-2 for metric scene geometry, and SAM 3D Body for each crew member as parameters of the MHR body model. No model was fine-tuned for microgravity. All numbers below are measured on held-out frames.
The simplest useful change is to stop sending pixels for the crew. In the parameter mode, SAM 3D Body runs on board and each crew member goes down as its MHR parameters in float16: per person per frame, plus once per person for body shape. The module goes down as a 320-pixel JPEG. The ground rebuilds meshes from the parameters exactly and rebuilds the module with MoGe-2. Figure 2 compares both modes on a 51-second held-out segment.
Captured at one frame per second; never sent in parameter mode.
SAM 3D Body and MoGe-2 each estimate distance from a single image, and they disagree. For every crew member we compared the mesh's visible surface with MoGe-2's depth at the same pixels, then scaled the mesh about the camera centre until they matched. Scaling about the camera keeps the mesh's outline in the image unchanged. Figure 3 shows the correction each person needed.
Could the station run only SAM 3D Body's image backbone and send its internal embedding, leaving the decoder on the ground? We patched the model at that point, compressed the embedding (float16, 8-bit, or PCA to k channels plus 8-bit, with the basis fitted once on training frames), and compared the ground's output with uncompressed SAM 3D Body on the same held-out frames and boxes.
| What is sent | Per person | Variance kept | Joint error | Pose erroraligned | Position error |
|---|
We trained our own encoder and decoder to reproduce both reference models from one latent. The encoder is a frozen DINOv2 ViT-S/14 with a small trained bottleneck (); the latent travels with a 32 × 18 colour thumbnail. The ground decoder predicts scene depth and a heatmap of crew centres with MHR parameters at each peak (CenterNet-style). Table 3 compares it with sending a small JPEG and running MoGe-2 on the ground.
| Downlink | Bytesper frame | Depth error | Shape error | ColourPSNR | Crew found | Joint error | Pose erroraligned |
|---|
A 2-billion-parameter vision-language model (Qwen3-VL-2B) reads each 10-second window and writes a short entry for controllers, sent as text (tier T0). Two design rules made the log usable. First, counting is delegated: the model is given the person detector's crew counts instead of counting across ten images itself. Second, it is told what it must not assume: crew float, and empty spacesuits are equipment. Its entries are visible in Figure 2. The log describes; it does not validate. Step validation belongs to a checker over measured state (N7).
Each direction below states the idea, why it matters for BAS, and the experiment that would confirm or reject it. Section 5 orders them into a plan.
Send the crew state every second, but the scene only when it changes. The ground and station run the same simple predictor (the last keyframe); the station sends a new keyframe when its view differs from what the ground would predict. Bandwidth then scales with activity rather than time.
The station's interior is designed, built and documented. Its geometry can live on the ground as a 3D model, registered to each camera once. The downlink then carries only what moves: crew, and hardware whose state changes.
SAM 3D Body has 840M parameters, too many for flight hardware. We distil it into a small student that predicts the same MHR parameters, trained with losses on parameters, on the rebuilt mesh, and on intermediate features. Training data comes from the teacher on real footage plus our synthetic generator, which renders crew at every orientation with exact labels.
V-JEPA 2 learns video representations by predicting the embeddings of future frames rather than their pixels. On board, the gap between what it predicts and what it then sees is a direct signal of the unexpected: a skipped step, a dropped object, an unusual posture. The same predictor, shared with the ground, turns the downlink into predictive coding: send only the residual.
VL-JEPA predicts the embedding of a text answer instead of generating tokens, and decodes to words only when needed. Protocol steps written as sentences ("open the glovebox port", "transfer sample A") become targets in the same embedding space. A new experiment then only needs its protocol text, not new training, and the log is decoded only when the recognised step changes.
In microgravity "upright" is meaningless. We express each crew member's pose in the rack's coordinate frame, not the camera's or gravity's, and train and test with bodies at every orientation from the synthetic generator.
A protocol checker compares the sequence of measured states (hand contacts, object positions, lid open or closed) against the expected procedure and raises explicit alerts. The language model narrates the checker's output; it never decides whether a step was valid.
E3 shows that generic compression of a large model's features fails. Codecs trained for the task, with a learned bottleneck placed where the task loss is measured, compress far better in published edge–cloud work. We would train one at the student's feature layer, with the ground-side decoder fine-tuned jointly.
| ID | Question | Success criterion | Status |
|---|---|---|---|
| E1 | Do body parameters plus a small image give the ground the same crew as full frames? | Identical meshes at ≥ 5× fewer bytes | done |
| E2 | Do body and scene models agree on scale, and can they be reconciled? | Mesh surface meets scene depth; plausible body heights | done |
| E3 | Is SAM 3D Body's internal embedding a good thing to send? | Accurate below full-frame size | done: no |
| E4 | Can one learned latent carry scene and crew? | Crew within 10 cm per joint at ≤ 5 KB | done: not yet |
| E5 | Can a small on-board model write a useful log? | Correct crew counts and hardware in each entry | done |
| E6 | How low can event-triggered keyframes take the downlink? (N1) | ≤ 2 KB/s with the same crew meshes | next |
| E7 | How small can the scene image be in parameter mode? | Shape error within 2 points of the 320-pixel image | next |
| E8 | Can a distilled student replace SAM 3D Body on board? (N3) | Aligned pose error ≤ 60 mm; ≤ 150 ms on 4 CPU cores | next |
| E9 | How much does body orientation hurt the reference model? (N6) | Error curve by orientation; fix recovers upright accuracy | planned |
| E10 | Step recognition and validation on the synthetic protocol (N5, N7) | ≥ 95% step accuracy; every injected mistake flagged | planned |
| E11 | Does predictive error flag protocol deviations? (N4) | Separates correct from wrong runs without labels | planned |