Computer Vision · Video Understanding · Multimodal AI · Vision-Language Models · Generative AI
Download CVI hold a PhD in Communication Technologies and Systems and focus on computer vision, video understanding, and multimodal AI. Over the last 6+ years I have built deep-learning systems for face recognition, sports video summarization, action recognition, and vision-language modeling — combining research and engineering across dataset creation, model training, large-scale experimentation, evaluation, and GPU-based deployment. I am especially interested in roles in AI research, applied science, research engineering, multimodal AI, video understanding, and generative AI.
Download My CVImage & video recognition, face analysis (ArcFace, RetinaFace), action recognition (SlowFast, VideoMAE), video summarization & highlight detection, Vision Transformers (ViT, Swin, DINOv2), 2D/3D CNNs (ResNet, I3D).
Vision-Language Models (CLIP, BLIP-2, LLaVA), Large Language Models (Llama, Mistral, Qwen), diffusion models (Stable Diffusion, SDXL, ControlNet), text-guided generation.
PyTorch, HuggingFace (Transformers, Diffusers), PyTorch Lightning, Weights & Biases, end-to-end ML pipelines, large-scale training & evaluation, Docker, Kubernetes, AWS/SageMaker, CUDA / GPU clusters.
AI Researcher / PhD Researcher
Universidad Politecnica de Madrid
Designed and evaluated deep-learning systems for computer vision and video understanding, including face recognition, sports-video summarization, and multimodal video analysis. Built end-to-end ML pipelines with CNNs, Vision Transformers, CLIP, and vision-language models, delivered applied R&D with Nokia and Airbus, and published at CVPR Workshops, IEEE Access, Scientific Reports, and AVSS.
Junior Programmer
Indra Sistemas S.A.
Developed automation scripts for Voice-over-IP communication systems in air-traffic-management environments. Configured and tested telecommunications equipment, routers, and switches, collaborating with engineering teams under Scrum methodology.
Data Coder & Trainee Programmer
DEYDE Calidad de Datos S.L.
Processed, validated, and maintained structured databases of municipalities, streets, and postal-address records. Contributed to rule-based expert systems for postal-address correction and coding, supporting data-quality and software-support workflows.
Ph.D. in Communication Technologies and Systems
Universidad Politecnica de Madrid, UPM
Thesis: Automatic Sports Video Summarization with Identity-Aware Highlight Selection. Grade: Cum Laude.
M.S. in Telecommunication Engineering
Universidad Politecnica de Madrid, UPM
Thesis: Automatic Highlight Detection in Martial Arts Tricking Videos. Grade: 10/10, with Honors (MatrÃcula de Honor).
B.S. in Telecommunication Systems Engineering
Universidad de Alcala de Henares, UAH
Thesis: Acoustic Bird Classification Using MFCC Feature Extraction. Grade: 9/10.
ViMoCLIP: Video Motion Cues for Animal Action Recognition
A novel approach that augments static CLIP representations with video motion cues for improved animal action recognition. Published at IEEE/CVF CVPR Workshops 2025.
Text-Guided Sports Highlights with CLIP
CLIP-based framework for automatic video summarization of soccer matches. Uses multimodal (text + image) neural networks for highlight detection.
Vision Transformers vs CNNs for Face Recognition
Comprehensive comparison between Vision Transformers and CNNs for face recognition tasks. Published in Nature Scientific Reports.
Automatic Highlight Detection in Martial Arts Tricking
Deep learning system for automatic highlight detection in martial arts tricking videos using 2D/3D CNNs, recurrent networks and Transformers.
UPM-GTI-Face: Face Detection Dataset
Dataset for evaluating the impact of distance and face masks on face detection and recognition systems. Presented at IEEE AVSS in Madrid.
Automatic Sports Video Summarization with Identity-Aware Highlight Selection
Doctoral dissertation presenting novel deep learning methods for automatic sports video summarization. Combines identity-aware techniques with highlight detection for personalized content generation.
Automatic Highlight Detection in Martial Arts Tricking Videos
Development of a deep learning strategy to automatically detect highlights in martial arts tricking videos. Grade: Distinction (10/10).
Acoustic Bird Classification Using MFCC Feature Extraction
Acoustic classification of bird species using sound feature extraction through MFCC (Mel-Frequency Cepstral Coefficients) parameters. Grade: Excellent (9/10).
Olympus: OpenClaw Agent System
Personal AI workspace built on OpenClaw and running 24/7 on a dedicated Mac Mini. It orchestrates specialized agents for coordination, planning, and coding, with persistent memory, Telegram-based control, task tracking, and GitHub-backed automation.
Paper Copilot: Research Paper Summarizer
Local AI agent that reads a research paper (PDF) and produces a structured, Notion-ready markdown summary including metadata, methodology, key results, reference analysis charts, and extracted figures.
Job Finder: AI Job Monitoring Agent
Local-first AI agent that crawls 16+ public job sources, normalizes postings, and scores relevance using hybrid ranking (rule-based, semantic embeddings, and LLM fit). Includes a Streamlit dashboard for review.
RoboMaster Tank Detection
Designed a university course assignment using RoboMaster tanks. Implemented localization networks in TensorFlow, object detection in PyTorch, and YOLOv8 for real-time detection.
Story & Image Generation
INDESIAhack: Weather Detection for Ferrovial
Led a team to build a Model of Experts (MoE) combining CLIP, Microsoft Azure, and ChatGPT API to assess road visibility from traffic cameras worldwide. Ran on AWS SageMaker.
Custom Components with TensorFlow
Organized workshops teaching custom TensorFlow components: loss functions, activations, initializers, regularizers, metrics, layers, models and training loops. Explored library internals and graph management.
Personal Website
Self-taught HTML, CSS, and JavaScript to build this personal portfolio. Learned web development from scratch, version control with Git, and deployment on GitHub Pages.
Kubernetes Cluster with 17 GPU Nodes
Built a Kubernetes cluster integrating 17 GPU-equipped computers. Configured NVIDIA support, native authentication, NFS storage, custom Docker profiles, and a minimalist JupyterHub interface for distributed neural network training.
Multi-GPU Workstation Assembly
Selected components and assembled multiple high-performance multi-GPU workstations for the research group. Handled system administration: OS installation, user management, package control, and driver updates.
How Transformer LLMs Work
A guided tour of the transformer architecture behind modern LLMs by Jay Alammar and Maarten Grootendorst: tokenization, embeddings, self-attention and transformer blocks, plus recent advances such as KV cache, multi-query / grouped-query attention, and Mixture of Experts.
Attention in Transformers: Concepts and Code in PyTorch
Josh Starmer's deep dive into the attention mechanism that powers LLMs: Query/Key/Value matrices, self-attention and masked self-attention, cross-attention, and multi-head attention — combining mathematical intuition with hands-on PyTorch implementations.
Generative AI with Large Language Models
AWS & DeepLearning.AI course covering the full generative-AI project lifecycle: transformer architecture, model pre-training and scaling laws, instruction fine-tuning, parameter-efficient methods (LoRA, soft prompts), evaluation, and reinforcement learning from human feedback (RLHF) for deployment.
How Diffusion Models Work
A hands-on short course by Sharon Zhou that builds diffusion models from scratch — the sampling process, noise-prediction neural networks, training, conditional and personalized generation, and techniques to speed up sampling — going beyond pre-built APIs.
Deep Learning Specialization
Andrew Ng's five-course specialization spanning neural networks and deep learning, network tuning and optimization, structuring ML projects, convolutional neural networks, and sequence models (RNNs, LSTMs, Transformers), with applications across computer vision, NLP, and speech.
I'm working on adding more content to this website.