Enabling voice
Voice uses OpenAI GPT-Live-1 and needs an OpenAI API key from a project with GPT-Live access. A deployment admin can enter one from Settings > Integrations > Voice, or a self-hosted operator can setR_VOICE_OPENAI_API_KEY; the environment variable is used when both exist.
Voice is opt-in: the deployment’s general OPENAI_API_KEY is not used, so
enabling OpenAI for task inference does not turn voice on. See the
Voice integration page for details.
The Voice integration settings include the GPT-Live voice selection and an
AI-generated audio preview. Existing selections are preserved. New and unset
connections default to Marin, one of OpenAI’s recommended voices.
When the voice key is not configured the voice button does not appear. The key
stays on the control plane. The browser sends its WebRTC connection offer to Roomote
and receives only the negotiated session answer; it never receives the API key.
Using voice
- Select the voice button in a composer: in an open Session, or on the home page and the New Session dialog. From the home page or dialog a new Session is created and the call starts inside it.
- Grant microphone access when the browser asks. A short rising tone confirms the call is open; a falling tone marks the end. A Call started marker appears in the Session.
- Talk to Roomote the way you would on a phone call. It acknowledges each utterance in a few words, hands it to the Fast Session, and reports the response out loud when it lands. Every utterance is sent to the Fast Session, where the selected model, tools, context, and safeguards handle the response.
- Speak at any time to interrupt. Roomote keeps listening while it speaks, and follow-ups go back through the same Fast Session. You can also type in the composer during the call.
- Use the in-call controls to mute your microphone, silence Roomote’s audio without muting yourself, or end the call. The button stays highlighted while the call is active, and a Call ended marker records its length.