Available for new projects · Prague, CZ
Fredy Rodriguez
Python Developer & Cloud Architect
I build intelligent systems that actually ship — LLMs running on servers I manage myself, cloud infrastructure tuned for every megabyte, and data pipelines that hold up in production.
- parameter model served locally
- 30B
- parameter model served locally
- live applications in production
- 3
- live applications in production
- self-managed cloud servers
- 2
- self-managed cloud servers
- monthly infrastructure bill
- $0
- monthly infrastructure bill
About
Lean, reliable systems — from the backend to bare metal
I specialize in efficient, scalable solutions for modern technical challenges. My work spans Python backends, cloud-native deployments and AI integrations, with a strong bias toward systems that run lean and stay up.
This site is the proof: it and a 30-billion-parameter language model both live on free-tier hardware that I provision, harden and maintain myself.
Python & backend
APIs, services and automation built for production with FastAPI, Streamlit and well-structured Python.
Cloud & infrastructure
Linux, NGINX, Docker and Terraform on Oracle Cloud — hardened, monitored and cost-efficient.
AI & LLM systems
Self-hosted inference with llama.cpp, LLM API integrations and LangChain workflows.
Data engineering
Pipelines and dashboards with Pandas, PostgreSQL, Redis and Plotly.
Toolkit
- Backend
- PythonFastAPIStreamlitREST APIsHugo
- AI / ML
- llama.cppGGUFPyTorchLangChain
- Infrastructure
- LinuxDockerNGINXOracle CloudTerraformCI/CD
- Data
- PandasPostgreSQLRedisPlotly
- IoT
- Raspberry PiArduinoMQTT
Projects
Live systems, not screenshots
Every project below is running right now on infrastructure I operate. Open one and try it.
Project 01 · LLM gateway
AI Chat Gateway
Authenticated chat interface for frontier LLMs. Switch between DeepSeek-V3, DeepSeek-R1 and Gemini 2.0 Flash depending on the task, with full conversation history.
Project 02 · Private inference
Self-Hosted LLM
Private chat running NVIDIA Nemotron 3.5 Lightning 30B-A3B entirely on my own Oracle Cloud Ampere ARM server — CPU-only inference with llama.cpp, served over HTTPS.
Project 03 · Data dashboard
Portfolio Analytics
Dashboard for tracking and analysing a Trading 212 portfolio: composition breakdowns, ETF vs. stock allocation, dividend history and performance comparisons.
Infrastructure
The machines behind it
No managed platforms, no vendor magic — two free-tier instances, provisioned and tuned by hand.
Node 01 · x86
Web & edge
- CPU
- 1 OCPU AMD EPYC
- Memory
- 1 GB
- Stack
- Hugo + NGINX (HTTP/2)
- TLS
- Let's Encrypt · auto-renew
Migrated from WordPress to a static Hugo build: memory use dropped from ~689 MB to a few MB, and every page is pre-compressed and served straight from disk.
Node 02 · ARM64
Inference
- CPU
- 4 cores Ampere Altra
- Memory
- 24 GB
- Runtime
- llama.cpp · ARM NEON
- Model
- Nemotron 3.5 Lightning 30B-A3B
A 30B mixture-of-experts model (~3B active per token) doing CPU-only inference — zero GPUs and zero external API calls. Every token is generated on hardware I control.
Beyond code
Languages
- SpanishNative
- EnglishFluent
- PortugueseFluent
- CzechCommunicative
Fun fact
I love coffee — and most mornings it comes from my family's own plantation in Colombia. Nothing beats starting the day with a cup from home.
Interests
Contact
Let's build something that runs well
Available for Python, cloud architecture and AI systems work — remote from Prague, or on-site across the EU.