androidinterview.com

Android System Design Interview Questions

How do voice and video calls work?

Tier: Less commonDifficulty: Hard

Explain how a calling app connects two people and carries their microphone audio and, optionally, camera video in real time. The app must establish the call, let the recipient accept or decline, carry the media and clean up when either person hangs up.

The problem

Start with a one to one call between phones on different networks. One path exchanges call setup messages, often called signalling. Another carries the live audio and video. Both need to work even when a direct connection between the phones is unavailable.

For example, Ana calls Ben. Ben's phone rings and he accepts. Both devices establish a media connection and begin exchanging audio. If Ben's network becomes weak, the app should adapt quality or show reconnecting instead of silently presenting a frozen call as healthy.

Cover call state, connection setup, media delivery and Android lifecycle integration. Group calls, recording and screen sharing are separate extensions. The transport explanation makes sense after these user visible stages are clear.

Starting the design

Voice and video calling is built on WebRTC, an open standard that gives browsers and mobile apps peer to peer real time media. WebRTC deliberately doesn't solve the whole problem on its own, so two pieces have to be built around it before any call connects, and on Android a third layer sits above it just to make the phone ring.

What I'd clarify first

  • Is this one to one calling only, or does group calling need to be in scope, since group calls usually need a media server rather than pure peer to peer.
  • Does a call have to survive the app being backgrounded, the screen locking, or the app not running at all when the call arrives, because that's a Telecom and foreground service question rather than a WebRTC one.
  • Does the app need to work through restrictive corporate or carrier networks, which changes how much you have to lean on relay infrastructure versus direct peer connections.
  • Is call quality adaptation to a poor network a requirement, or is a fixed quality level acceptable.

The two problems WebRTC doesn't solve

  • Signaling. Before two devices can exchange media they have to exchange setup information, session descriptions, network candidates, who's calling whom. WebRTC intentionally leaves this out, so you build it, typically over the same WebSocket infrastructure a chat feature already has, since it's just another kind of real time message passing.
  • NAT traversal. Most phones sit behind a carrier or router NAT with no public address a peer could dial directly. STUN and TURN solve it. STUN asks a public server what your actual public address is, which is often enough for two peers to reach each other directly. When it isn't, common on strict corporate networks, a TURN server relays the media between the two peers, which costs latency and real bandwidth because the traffic is no longer taking the direct path.

Encryption is worth naming as the one thing WebRTC does solve for you. Once a path is picked, a DTLS handshake runs over that same path and derives the keys for SRTP, which encrypts every media packet. It's mandatory in the spec, there's no unencrypted mode to fall into by mistake. What that protects is the hop between the two endpoints, so on a group call the media server sits inside the boundary and can see the media unless you add end to end encryption on top of it.

How a call connects

  1. The caller creates an offer, a description of what it can send and receive, and pushes it over the signaling channel.
  2. The callee's device rings, and on accept it sends back an answer in the same form.
  3. Both sides gather ICE candidates at once, the local address, the public address STUN reports, and a TURN relay address, and trickle each one across the signaling channel as it's found rather than waiting for the whole set.
  4. Both sides run connectivity checks on the candidate pairs concurrently and settle on the best pair that works, preferring a direct path over a relayed one.
  5. DTLS handshakes over that path and hands SRTP its keys.
  6. Audio and video flow on the media path, and the signaling channel drops back to carrying call control only, mute state, hang up, and so on.

What Android adds

  • Ringing when the app isn't running. The invite arrives as a high priority FCM data message, which can improve delivery urgency in Doze but does not guarantee immediate arrival. Your handler wakes, posts the incoming call UI immediately, and only then starts any signaling. A normal priority message can be held until the next maintenance window, which for a call is the same as never.
  • The incoming call UI. Notification.CallStyle gives you the ringing treatment with answer and decline actions the system understands. To take over a locked screen you attach a full screen intent, and from Android 14 the USE_FULL_SCREEN_INTENT permission is granted by default only to calling and alarm apps, so check canUseFullScreenIntent and degrade to a heads up notification rather than assuming you have it.
  • Being a real call. Register with the Telecom framework, these days through androidx.core.telecom and its CallsManager rather than writing a ConnectionService yourself, and declare MANAGE_OWN_CALLS. That's what makes an incoming carrier call put yours on hold instead of both playing at once, and what makes the answer button on a Bluetooth headset, a watch or Android Auto do anything.
  • Staying alive for the duration. Use the supported Telecom lifecycle and applicable foreground service types for ongoing calls and capture. A custom microphone or camera service requires the matching permissions and must satisfy while in use and background start restrictions. A foreground service does not guarantee an immortal process. Coordinate audio focus and routing through the calling integration instead of adding competing policies.

How I'd build this on Android

Calling use casesThe system boundary Calling holds the use cases Start or answer call, Mute or switch camera, Change audio route, End call. Caller or callee takes part in Start or answer call, Mute or switch camera, Change audio route, End call.
Calling use cases, a use case diagram for Calling
Android call control and mediaCall route and ViewModel to Call session owner, observe and control. Call session owner to Signaling repository, call setup. Call session owner to WebRTC media engine, capture and render. Signaling repository to Signaling backend, control messages. WebRTC media engine to Peer, TURN or media server, media packets.observe and controlcall setupcapture and rendercontrol messagesmedia packetsCall route and ViewModelControls, state and video surfacesCall session ownerTelecom and lifecycle coordinationSignaling repositoryInvites, offers and call stateWebRTC media engineTracks, ICE and encrypted mediaSignaling backendIdentity and session routingPeer, TURN or media serverReal time media path
Android call control and media
Signaling and media are distinct paths. Rotating the UI does not create a second call session.

I'd give the call a session owner that can outlive an Activity. CallViewModel observes it and exposes mute, camera and hang up actions. A repository handles signaling, the messages that set up and end the call. A media adapter manages WebRTC tracks and connections. Room can save call history, but audio and video packets never pass through a DAO.

interface CallSession {
    val state: StateFlow<CallState>
    suspend fun setMuted(muted: Boolean)
    suspend fun end()
}
data class CallState(val callId: String, val phase: String, val muted: Boolean)

class CallViewModel(private val session: CallSession) : ViewModel() {
    val uiState = session.state
    fun mute(muted: Boolean) = viewModelScope.launch { session.setMuted(muted) }
    fun hangUp() = viewModelScope.launch { session.end() }
}

The Compose route collects state with lifecycle awareness and releases its video renderer when the UI is removed. Rotating the screen removes a surface without ending the call. When the call ends or fails permanently, the session releases tracks, camera, microphone and the peer connection. Telecom coordinates calls and audio routing. I'd follow that integration's lifecycle instead of adding another audio focus policy that can conflict with it.

Microphone and camera permission requests belong in the UI before capture starts. An FCM invite can be late or downgraded, even at high priority. Before showing it, the app checks the call ID, current status and expiry. It also follows Android rules for full screen intents, notifications and foreground service starts. WorkManager cannot carry live media or guarantee an incoming call rings at an exact time.

I'd inject signaling, the media engine, session store and clock. A fake session tests ringing, denied permission, reconnect, remote hang up and rotation. Device tests cover a carrier call interrupting the app, Bluetooth route changes and losing the camera. I'd measure call setup time, packet loss, round trip delay and unexpected endings. See building a Telecom calling app.

Tradeoffs I'd call out

  • Peer to peer vs a media server. Direct peer to peer gives the lowest latency and costs the platform nothing in relay bandwidth, but every peer sending its stream to every other peer stops scaling past three or four participants. Past that you need a media server, and there are two kinds. An SFU takes one stream from each participant and forwards each one on unchanged, cheap on server CPU and heavier on the receiver's bandwidth and decode load. An MCU decodes everything and mixes it into a single stream per participant, the reverse trade, expensive on the server and easy on the client. SFU is the default answer now, MCU shows up where clients are weak or the output has to be one stream.
  • Preferring the direct path vs going straight to TURN. This isn't a serial timeout, which is the thing people get wrong. ICE gathers host, STUN and TURN candidates at the same time and checks the pairs concurrently, so the cost of preferring direct is that the calls which end up needing the relay connect slightly later, not that every call waits for a direct attempt to fail first. Forcing TURN on everything is more predictable, at the price of paying relay bandwidth for calls that would have been free.
  • Fixed call quality vs adaptive bitrate. A fixed resolution and bitrate is simpler to reason about, but on a degrading connection it freezes or drops. Adaptive bitrate, which WebRTC does natively, walks resolution and frame rate down smoothly as bandwidth falls. Opus on the audio side already adapts its own rate, and on video you're picking between VP8 and H.264, which every endpoint has, and VP9 or AV1, which cost more CPU for a better picture at the same bitrate.

What breaks at scale and on a poor connection

At scale the cost center isn't signaling, it's TURN relay bandwidth. Every call that can't connect peer to peer costs the platform real, ongoing bandwidth for its whole duration, which is why call quality and network diagnostics matter for capacity planning and not just for the person on the call. On a poor connection, adaptive bitrate carries you a long way, and when it isn't enough the right move is an automatic drop from video to audio only. A call that degrades to audio and keeps going is a far better outcome than one that holds out for video quality and freezes.

Read more Build a calling app (opens in a new tab)Foreground service types (opens in a new tab)

Watch