RFD 1095: A voice XR path, beside Task Manager, same backend

Problem

Task Manager already drives 3DAIGC-API through REST, but a headset user with hands full, or wanting a spoken “make a 3D model of this,” has no voice-and-camera path to the same inference backend.

Decision

Wire NVIDIA’s xr-ai stack, running on the CUDA GPU host, to 3DAIGC-API through an HTTP MCP adapter, not a second inference backend. Any WebXR headset’s camera and microphone reach the XR Media Hub (port 8088, LiveKit for WebRTC on 7880-7882), which runs speech-to-text, a vision-language model, then text-to-speech; a voice-agent sample (3daigc-vlm-example) calls the same 3DAIGC-API mesh jobs through MCP tools (upload_image, image_to_textured_mesh, wait_for_job). When a router blocks the headset from reaching the GPU host directly, scripts/xr-spark-hub-proxy.mjs on the editor workstation forwards :8443 to the GPU host’s own :8088. RFD 1086’s Surface and DGX Spark are this team’s own reference pair, tested with Galaxy XR; RFD 1119 gives the general requirement.

See DETAILS.md for the full port table, the start and monitor scripts, the MCP tool flow, and the troubleshooting table.

Details

The measurements and the retractions