Combining Local Processing with Cloud Serverless for Optimal AI Inference
Artificial intelligence development teams often face a significant decision regarding inference: whether to self host or use cloud services. Self hosting involves managing GPUs,...
Artificial intelligence development teams often face a significant decision regarding inference: whether to self-host or use cloud services. Self-hosting involves managing GPUs, which can be costly when underutilized. Alternatively, cloud APIs offer quick starts but come with data security concerns and model limitations.
Rather than choosing one approach exclusively, a hybrid strategy can be more effective. This involves splitting the workload, keeping some inference tasks on local hardware while leveraging serverless cloud services for others. Deciding the division point is crucial and should be based on specific workload characteristics.
Example: Speech-to-English Translation
Consider a tool designed for speech-to-English translation. It performs automatic speech recognition (ASR) on local hardware and uses a serverless platform for translation. The tool is available as an open-source project on GitHub. This setup exemplifies a method to determine which tasks should remain local and which should be cloud-based.
System Architecture
[ LOCAL HARDWARE ] [ SERVERLESS CLOUD ]
audio input → ASR model → transcript → translation → English text
(mic / file) on-device (MPS/CPU) (text only)
The ASR step is managed locally, ensuring that audio data doesn't leave the user's device. Only the text transcript is transmitted for translation, which is handled serverlessly. This approach balances privacy, cost, maintenance, and capability access.
Deciding Local vs. Serverless Inference
To determine whether an inference step should be local or serverless, consider four main factors:
- Privacy/Data Residency: Sensitive data should remain local, while sanitized data can be processed in the cloud.
- Cost Shape: Frequent tasks benefit from local execution, whereas bursty tasks are well-suited for pay-per-use cloud services.
- Maintenance Burden: Small models can be managed locally, while larger models are better suited for serverless hosting.
- Capability Access: Utilize local resources when available, and opt for hosted models when provisioning is complex.
Privacy Considerations
Audio data is highly sensitive, containing personal information like voiceprints. By processing audio locally and sending only text transcripts, exposure is reduced. This approach significantly decreases the amount of data transmitted, enhancing privacy.
Cost Implications
Running ASR locally is cost-effective due to its high frequency. In contrast, translation is less frequent and benefits from a pay-per-use model. For example, serverless translation costs can be minimal per translation instance.
Maintenance and Capability
The ASR model is compact, suitable for consumer-grade hardware, while the translation model's complexity is managed by the serverless provider. This reduces operational overhead and simplifies access to advanced capabilities without extensive setup.
Addressing Constraints
Initially, the plan was to handle both transcription and translation in the cloud. However, testing revealed limitations with audio payloads. By moving ASR locally, the architecture not only overcame these constraints but also improved privacy and cost efficiency.
Verification and Code Example
The following code exemplifies the serverless call for translation:
from openai import OpenAI
client = OpenAI(
base_url="https://inference.do-ai.run/v1",
api_key=os.environ["MODEL_ACCESS_KEY"],
)
response = client.chat.completions.create(
model=os.environ.get("DO_MODEL", "nemotron-3-nano-omni"),
messages=[
{"role": "system", "content": "Translate the user's text to English."},
{"role": "user", "content": transcript},
],
timeout=float(os.environ.get("DO_TIMEOUT_SECONDS", "90")),
)
This straightforward integration highlights the simplicity of bridging local and serverless components. The code managing the handoff between local ASR and cloud translation is minimal, emphasizing the ease of maintaining such a system.
Generalizing the Hybrid Pattern
This hybrid approach extends beyond translation. It can be applied to any task involving sensitive and high-frequency processing alongside resource-intensive operations. Examples include:
- Local redaction followed by serverless summarization.
- Local embedding generation with serverless reasoning.
- On-device data capture with serverless enrichment.
In summary, the strategy is to keep tasks that benefit from local execution and privacy on-device, while renting cloud services for more demanding operations as needed. This flexibility allows for incremental adoption of serverless technology, balancing control and efficiency.