TL;DR #
We built a containerized disposable Browser as a Service infrastructure where we stream logs from Docker to Python with specific tokens to strike a balance between speed and stability.
Context #
Almost two years ago, early in 2025 we set out to solve OSINT for organizations. We had seen the problems with shadow IT in a large organization so one of the first problems to solve seemed obvious - users need to be able to access unregulated web with minimal overhead. So Glazer Browser was born.
Glazer Browser, boiled down to its core, is a secure containerized environment that uses an anonymous network gateway, sends a video and audio feed of its desktop to the end user and receives mouse and keyboard events from them. There are a lot more nuances to the product but that’s the core. There’s also a realtime data collection and enrichment side to it but we’ll be talking about that in another post.
For building the sandboxed throwaway environment we first needed a virtualization layer. We try to be fairly pragmatic so one aspect of our decision making is the existing knowledge in the team. So we landed on Docker. Our end of the orchestration is written in Python.
Problem #
For building a truly isolated setup we need to spin up a sophisticated environment that guards the network traffic, handles storage, ACL, truly isolates the users and does this in a reasonable timeframe. We started off about 8-10 seconds per sandbox and over a year we’ve reduced that number to about 2 seconds which we are quite proud of. Of course, in the modern web it still seems like a long time so our work polishing that number goes on.
I’ll simplify our environment to two containers - sidecar and sandbox. Sidecar handles the network, sandbox the actual environment for end user. Our first attempt was a truly parallel startup process which was the fastest setup we’ve ever had. Environment was ready for the end user, at best in about 1.5-2.5 seconds. Problem was the error rate. In reality, that setup was reliant on race conditions all over the place. Hard to reproduce network errors and retry loops based on arbitrary numbers led to a very unstable user experience.
Before the initial launch, we migrated to a sequential startup process which was slow but almost bulletproof. But it was obvious for us that we need to find a system that can provide an environment for the user much faster.
Solution #
Our first instinct was to hold a pool of environments ready for the end user and then allocate them on-demand. As we researched this, we found multiple problems that we did not have resources to overcome. For example the wide combination of egress types we provide - different VPNs with their exit countries and Tor. We did not want to hide the country selection behind a region just so we could simplify the pre-generated pool. And without the up-and-running networking stack pooling didn’t really provide any significant boost for our setup.
Instead we took a hard look at the initial parallel setup, analyzed the points of failure there and worked out a system where each building block can spin up individually but also wait for the appropriate signal from its peers. At first, we failed because we treated “container is running” as “container is ready”. For our sidecar, that was dangerous as the proxy might be running in there but not actually be ready to carry traffic. So naturally, we inverted it, instead of guessing when the container might be ready we made each building block announce its readiness to the orchestrator.
So for the Tor sidecar it looks something like this
# start.sh
# Sidecar entrypoint. Same shape for every egress backend we ship.
# 1. Start the slowest thing first
start_tunnel & # tor client for example
# 2. Do the other things
setup_dns_routing
# eth1 is the isolated network
until ip link show eth1 > /dev/null 2>&1; do sleep 0.1; done
# Pull up the proxy server
proxy-server --config /etc/proxy.conf &
PROXY_PID=$!
until listening_on "$PROXY_PORT"; do sleep 0.1; done
# 3. Only now block on the long tunnel startup.
# Ready means tor can actually carry traffic, not merely that the port is bound.
if ! timeout "${TUNNEL_TIMEOUT:-25}" wait_for_tunnel_established; then
echo "tunnel failed to establish within ${TUNNEL_TIMEOUT:-25}s"
exit 1
fi
# And announce the readiness - more about this in a moment
echo "SIDECAR_READY"
# 4. The proxy process is the container's lifetime.
wait $PROXY_PIDThe script is built around a simple principle - run the slow thing and use that latency to do the remaining work. Establishing the tunnel dominates everything else so it’s launched immediately in the startup script. Then we set up the routing, waiting for the isolated network, starting the proxy and confirming its readiness. All of this is free since we have to wait for the tunnel that’s negotiating anyway. And then the container announces itself with a pre-configured token, “SIDECAR_READY” in the example above, to the orchestrator. Finally we wait on the PROXY_PID, which keeps the container alive for exactly as long as the proxy is alive. If the proxy dies, the container dies with it. That turns out to matter on the other side.
Reading the signal #
So the container announces itself on stdout. Nothing else was really an option for us here. The sidecar’s whole job is to be a network boundary, so a health endpoint would mean opening a route into the thing we’re trying to lock down, and it would have to be open before the network is even up - which is exactly the window we’re trying to watch.
On our end, docker-py gives us that stream as a generator, so in principle it’s a for loop:
startup_timeout = 30.0
stream_until = int(time.time() + startup_timeout + 5) # +5s grace for asyncio cleanup
request_id = get_request_id()
thread_event = Event()
def _wait_for_sidecar_ready() -> bool:
set_request_id(request_id)
try:
for line in sidecar.logs(stream=True, follow=True, until=stream_until):
if thread_event.is_set():
break
if line.strip() == b'SIDECAR_READY':
return True
# Stream ended and we never saw the token, so the container exited.
# Grab the tail while it's still around.
logger.warning(
'Sidecar log stream ended without SIDECAR_READY',
extra={'sidecar_id': sidecar.id, 'sidecar_logs': _safe_container_logs(sidecar)},
)
return False
except APIError as e:
if _is_container_gone(e):
# Teardown or the user cancelled. Not an error.
return False
logger.exception('Docker API error while streaming sidecar logs')
return FalseIt’s a short function but pretty much every line in it is there because something went wrong first.
until is an absolute timestamp, not a duration. It tells the Docker daemon when to close the stream on its side, whether or not anyone is still reading. Without it a client that dies without cleaning up leaves a follow stream open forever. We give it 5 seconds more than our own timeout so that our timeout always fires first - otherwise the stream just closes under us and we report it as “the container exited without signalling”.
The call blocks, so it can’t sit on the event loop. It goes into a thread through run_in_executor, and that means we can’t cancel it. Calling .cancel() on a future that already started does nothing at all. So we check a threading.Event once per line instead. It’s cooperative, and it’s the only option we have.
set_request_id(request_id) at the top of the thread is easy to miss. It’s our own helper - a contextvar our logger reads, so every line from one request carries the same id. contextvars don’t follow you into executor threads though, so without setting it again by hand every log line from the most failure-prone few seconds of the request loses its correlation id. Which is of course exactly when you want it.
And a stream that ended is not the same thing as a stream that timed out. If the stream ends and we never saw the token, the process died and we have something to look at, so we pull the last lines immediately, before teardown removes the container. If we hit the deadline instead, something is just slow. We used to collapse those two into one “startup failed” and it’s a good way to end up with an error rate you can’t explain. Same for an APIError that turns out to mean “no such container” - that’s usually teardown or the user beating us to it, and logging it as a failure is just noise.
The sandbox is a different story. Here we just poll from outside - the token mechanism wouldn’t buy us enough to justify the complexity:
def _wait_for_sandbox_ready() -> bool:
set_request_id(request_id)
while not thread_event.is_set():
result = sandbox_container.exec_run('pidof chromium')
if result.exit_code == 0:
return True
time.sleep(0.1)Then we run both of them against a single deadline, and the polling tick doubles as our progress update to the UI:
sidecar_future = loop.run_in_executor(None, _wait_for_sidecar_ready)
sandbox_future = loop.run_in_executor(None, _wait_for_sandbox_ready)
deadline = loop.time() + startup_timeout
while True:
remaining = deadline - loop.time()
if remaining <= 0:
thread_event.set()
for _ in range(20): # give the threads up to 2s to drain
if sidecar_future.done() and sandbox_future.done():
break
await asyncio.sleep(0.1)
raise SecureNetworkCreationTimedOut()
done, pending = await asyncio.wait(
{sidecar_future, sandbox_future},
timeout=min(0.25, remaining),
)
if pending:
yield _update_progress(remaining)
else:
breakThe deadline is set once, up front, and not reset on every wait. If you reset it, two components that are both almost ready can keep extending the total wait between them, and that’s a fun one to find in production.
On timeout we set the stop flag and then actually wait for the threads to finish before raising. Both of them are holding an open Docker stream, so raising right away leaks the thread and the connection. Under load that adds up fast.
Where we landed #
The progress numbers are partly made up - every tick with work still pending just nudges the bar along. Users usually see the progress bar inching a little and then leaping to the next step as soon as containers have said they are ready. Before this work our startup messages were really just us narrating our own API calls. For example, “creating the network” meant we had asked Docker for a network, not that there was one.
The numbers ended up in a nice place. The original fully parallel setup could do 1.5 to 2.5 seconds on a good day, but it was broken. The sequential version we shipped to buy stability cost us 8 to 10. Rebuilding the startup around declared readiness got us to about 4.5 on a cold boot, and caching the whole Tor data directory so the tunnel doesn’t have to bootstrap from scratch took that down to about 2.
The averages are the boring part though. Without the cache a cold boot landed anywhere between 3.4 and 6.6 seconds and we had no way of knowing which one a given user was going to get. With it, every run comes in between 1.9 and 2.1. That’s the same trade as everything else in this post - what we actually wanted wasn’t the smaller number, it was the narrow range.
Every building block starts as early as it can, does its slow work in parallel with everyone else’s, and nobody moves on until the thing they depend on has said it’s ready. It’s not all clean, though. The SIDECAR_READY token is compared inline, and it’s really an interface across several repos held together by an echo on one side and an == on the other.
Next up we’ll open the sidecar, which in this post was just a black box that prints a word. Several very different egress backends behind one contract, the data directory cache that bought us those two and a half seconds, and what fail-closed has to mean when a container dies mid-session.