Skip to content

Available for new projects · Prague, CZ

Fredy Rodriguez

Python Developer & Cloud Architect

I build intelligent systems that actually ship — LLMs running on servers I manage myself, cloud infrastructure tuned for every megabyte, and data pipelines that hold up in production.

parameter model served locally
30B
parameter model served locally
live applications in production
3
live applications in production
self-managed cloud servers
2
self-managed cloud servers
monthly infrastructure bill
$0
monthly infrastructure bill

About

Lean, reliable systems — from the backend to bare metal

I specialize in efficient, scalable solutions for modern technical challenges. My work spans Python backends, cloud-native deployments and AI integrations, with a strong bias toward systems that run lean and stay up.

This site is the proof: it and a 30-billion-parameter language model both live on free-tier hardware that I provision, harden and maintain myself.

Full experience on LinkedIn

Python & backend

APIs, services and automation built for production with FastAPI, Streamlit and well-structured Python.

Cloud & infrastructure

Linux, NGINX, Docker and Terraform on Oracle Cloud — hardened, monitored and cost-efficient.

AI & LLM systems

Self-hosted inference with llama.cpp, LLM API integrations and LangChain workflows.

Data engineering

Pipelines and dashboards with Pandas, PostgreSQL, Redis and Plotly.

Toolkit

Backend
PythonFastAPIStreamlitREST APIsHugo
AI / ML
llama.cppGGUFPyTorchLangChain
Infrastructure
LinuxDockerNGINXOracle CloudTerraformCI/CD
Data
PandasPostgreSQLRedisPlotly
IoT
Raspberry PiArduinoMQTT

Infrastructure

The machines behind it

No managed platforms, no vendor magic — two free-tier instances, provisioned and tuned by hand.

Node 01 · x86

Web & edge

CPU
1 OCPU AMD EPYC
Memory
1 GB
Stack
Hugo + NGINX (HTTP/2)
TLS
Let's Encrypt · auto-renew

Migrated from WordPress to a static Hugo build: memory use dropped from ~689 MB to a few MB, and every page is pre-compressed and served straight from disk.

Node 02 · ARM64

Inference

CPU
4 cores Ampere Altra
Memory
24 GB
Runtime
llama.cpp · ARM NEON
Model
Nemotron 3.5 Lightning 30B-A3B

A 30B mixture-of-experts model (~3B active per token) doing CPU-only inference — zero GPUs and zero external API calls. Every token is generated on hardware I control.

Beyond code

Languages

  • SpanishNative
  • EnglishFluent
  • PortugueseFluent
  • CzechCommunicative

Fun fact

I love coffee — and most mornings it comes from my family's own plantation in Colombia. Nothing beats starting the day with a cup from home.

Interests

TennisBadmintonElectronics & IoTAI research

Contact

Let's build something that runs well

Available for Python, cloud architecture and AI systems work — remote from Prague, or on-site across the EU.