Comunidad de inteligencia artificial en español. Noticias diarias, catálogo de herramientas open source y SaaS votado por la comunidad, y guías prácticas.

España
Pinned Tweet
Cada día en esta cuenta: noticias de IA en español, herramientas nuevas del catálogo (ya van más de 100, open source y SaaS) y a las 19:00 un hilo con lo mejor de la jornada. La comunidad vota, propone y comenta. Nosotros lo contamos claro y sin humo.
2
1
18
7,166
Cloudflare ha liberado Clef y Clef-flash: modelos pequeños que devuelven una decisión tipada con probabilidades. Su equipo de inteligencia de amenazas clasifica una web en 2,2 s con Clef frente a 4,7 s con gpt-oss-120b. ¿Qué decisión de tu agente moverías a un modelo así?
Yesterday AWS. Today Cloudflare. The decision-model land grab is on. Cloudflare just open-sourced Clef and Clef-flash. Small models that do one job. Take inputs, return a typed decision with probabilities. Which team picks up this ticket. Is this domain phishing. Should the agent hand over to a human. They claim the top of the Jev Decision Index and they're already selling an RL fine-tuning platform so you can bend one to your own categories. The number that got me: their threat intel team classifies a website in 2.2s with Clef, against 4.7s for gpt-oss-120b, and the small model returns richer output. So people call a large open model to make a routing decision and pay in latency for nothing. Which gets at the token economics. An agent run is mostly decisions, not reasoning. Where does this go next. Which tool now. Is this output good enough to accept. We've burned frontier tokens on those calls because nothing better existed. Now three vendors in a month have built exactly that thing. Jev, AWS's Strands Decider, and Clef. Some scepticism, though. The category is three weeks old and already has a leaderboard, a "we're number one" blog post, and a paid RL product attached. That's very 2026. But the core idea holds. Stop paying premium prices for glorified coin flips and save the big models for the calls that actually need thinking. Where would a decision model slot into your stack? Curious how many of your agent steps are basically routing.
55
Lo util es el camino offline: si el prefijo ya esta en tu historial sugiere sin llamar a nada, y solo pregunta al modelo cuando escribes una descripcion. 0,7-0,9 s por pulsacion, casi todo red. ¿Te vale esa espera por acertar el comando a la primera?
someone hooked Jev into zsh autocomplete. now your terminal knows what command you're about to type not pattern matching. not frequency ranking. an actual probability score. you type git ch. Jev reads your last 100 commands and returns: git checkout main [0.970] press → to accept. done. the part that makes it interesting: two modes running simultaneously. prefix mode - if any history entry literally starts with what you typed, Jev ranks those. fuzzy mode - if nothing matches literally, Jev reads all 100 entries and figures out what you mean. last 5 commits → git log --oneline -5 [1.000] you typed a description. terminal returned the command. one Jev request. two questions in parallel: Choice: which command is the user most likely completing? Noul: does any command in history actually complete this input? both signals run because they fail in different places. the Choice spreads out on nonsense. the Noul under-fires on tiny histories. together they cover each other's blind spots. latency: 0.7-0.9s per keystroke. almost all of it API time. Node startup and history parsing add 0.1s. cost per suggestion: fractions of a cent. Jev output is free. the offline path still works - a single prefix match gets suggested immediately without any request at all. a terminal that understands what you mean instead of what you type.
1
98
Las seis pruebas son las mismas para los siete routers, asi que las diferencias de precio y latencia se pueden comparar de verdad. Enrutar por tarea suele ahorrar mas que cambiar de modelo. Tu enrutas por tarea o tiras siempre del mismo modelo?
OpenRouter ha puesto 7 routers uno al lado del otro. Calidad, velocidad y precio. Seis pruebas. Jev Router está en la lista. Hasta ahora elegías modelo a ojo. Aquí ves cuál enruta mejor y cuál te sale más barato. openrouter.ai/benchmarks/rou…
1
100
Que un mod sea una función de TypeScript dentro de un plugin significa que puedes reescribir el prompt antes de que salga, sin tocar el modelo. El /diff de siempre ya funciona así. ¿Qué paso de tu flujo reescribirías tú?
Anthropic shipped Claude Code mods: small TypeScript functions that can rewrite prompts, tool calls, or the UI, and they ship inside plugins. Some built-ins like /diff are mods now. Feels like Claude Code just grew a plugin API for its own brain
2
95
Con cursor por clave indexada la pagina 1000 cuesta lo mismo que la 1: WHERE id menor que el ultimo visto, ORDER BY id DESC, LIMIT 20. Pierdes el salto a una pagina concreta, que casi nadie usa. Tu listado va con OFFSET o con cursor?
Tu paginación con LIMIT/OFFSET funciona perfecto en desarrollo y se cae sola en producción. El problema no se ve hasta que la tabla crece. Cuando pides la página 1000 con OFFSET 50000, la base de datos no salta a la fila 50000. Lee las 50000 anteriores, las ordena, las descarta una por una y recién ahí te devuelve las 20 que querías. Mientras más avanzas, más lento: la página 1 vuela, la página 5000 tarda segundos. Lo peor es la corrupción silenciosa. Si alguien inserta o borra filas entre página y página, OFFSET se desincroniza. Ves registros repetidos o te saltas otros, porque la posición numérica cambió debajo de ti. La alternativa es keyset pagination. En vez de "salta 50000 filas", le dices "dame las 20 siguientes DESPUÉS de este id": Lento (OFFSET): SELECT * FROM posts ORDER BY id DESC LIMIT 20 OFFSET 50000; Rápido (keyset): SELECT * FROM posts WHERE id < :ultimo_id_visto ORDER BY id DESC LIMIT 20; El keyset usa el índice directo: salta al id y lee 20. Es igual de rápido en la página 1 que en la 50000, porque nunca toca lo que ya pasó. Reglas: - Ordena por columna indexada y única (id, o created_at + id como desempate). - Guarda el último valor visto, no el número de página. - Para paginador con números clickeables en tablas chicas, OFFSET sirve. Para feeds, scroll infinito y APIs que crecen, keyset. OFFSET no está roto, está mal usado. Es cómodo para un admin con 300 filas. Para cualquier cosa que escale, es deuda técnica que se cobra sola.
1
1
164
Lo util aqui es que los requisitos son el texto literal de la pagina oficial, no un resumen generado. Y si no lo tiene claro, pregunta en vez de rellenar. Que tramite de tu pais montarias asi primero?
Peru has +32k government pages on gob.pe So we built "Hola, Perú", describe your situation, get the right official procedure 🇵🇪 → the requirements are the official page text → when it isn't sure, it asks instead of guessing → no account, open source
1
116
Que el 57,4% de los dominios citados salga solo para una marca explica que no haya una lista de webs donde colocarse: cada sector tiene la suya. Y tu propia web solo se lleva el 4,8% de las citas sobre tu negocio. ¿Ya has mirado quién te cita en ChatGPT?
📊 ESTUDIO de GEO: ¿Cuales son los 50 dominios que la IA cita para más marcas en España? Lo que mueve a ChatGPT no mueve a Google. Puedes salir bien en un motor y ser invisible en otro sin enterarte. Estudio de Mencoro.com: 176.332 citas en 23.466 respuestas de IA, 42 marcas en España, 8 sectores. Del 1 de julio al 30 de septiembre de 2026, en ChatGPT, AI Overviews, Modo IA de Google y Perplexity. Lo que sale: ➡️ Tu web se lleva el 4,8% de las citas en las respuestas sobre tu propio negocio. En ChatGPT, el 3,5%. ➡️ El 72,6% va a webs de empresas: tus competidores, tiendas, clínicas, fabricantes. ➡️ El 57,4% va a dominios que solo aparecen para una marca. No existe una lista universal de webs donde hay que estar. ➡️ Google tira de redes: el 97,6% de las citas a Instagram sale de AI Overviews y Modo IA. ➡️ ChatGPT no citó TikTok ni una vez. Pero suyas son el 84,1% de las citas a El País y el 94,4% de las del BOE. Para quien piense que el SEO ha muerto: las páginas que compiten por esas citas son fichas de producto, de servicio y de precios. SEO de toda la vida, ahora leído por cuatro motores. ¿Qué deberías revisar esta semana? ➡️ Mide cada motor por separado. Una media te esconde dónde no existes. ➡️ Mira qué páginas de tus competidores se citan en tus preguntas de compra. ➡️ Completa tus perfiles en los directorios y marketplaces de tu sector. ➡️ Si tu comprador pregunta en ChatGPT, pesa la prensa nacional. Si busca en Google, pesan tus redes. ¿En qué motor crees que tu marca está peor? Estudio completo y CSV descargable: mencoro.com/es/estudios/domi…
2
130
Este AGENTS.md hace que Claude Code y Codex trabajen más como un ingeniero senior. Codex lo lee de forma nativa y las versiones recientes de Claude Code también pueden utilizarlo cuando no existe un CLAUDE.md. Incluye 6 reglas bastante simples: → Planifica antes de tocar código → Haz el cambio mínimo necesario → Divide trabajo entre subagentes → Hazte responsable del bug → Verifica antes de darlo por terminado → Guarda cada corrección importante Lo interesante es que el archivo mejora con el tiempo, cada corrección hace que el agente se adapte mejor a tu forma de trabajar. Copia el archivo en tu AGENTS.md 👇
15
30
159
11,839
Lo interesante de que todo sea plugin: mandas lo rutinario a un modelo local y dejas la API solo para lo difícil, sin cambiar de herramienta. ¿Qué modelo local te está funcionando bien para el día a día?
¡DeepSeek Harness disponible en Windows y macOS! Alternativa open source de la app de Claude y ChatGPT para trabajar en tu ordenador. ✓ Todo es un plugin. 100% configurable ✓ Usa cualquier modelo de IA (local o API) → deepseek.com/harness/
2
81
A claude sonnet 5.5 le tomó 7 minutos, mucho menos que a GPT-6 astra, el rendimiento de salida de sonnet es mejor así que gana sonnet. (SARCASMO)
I gave Claude Sonnet 5.5 and GPT-6 Astra the same challenge "Build what you think you look like"
3
1
190
26,293
Tengo ambos planes de 200$ de claude y de chatgpt. Mi experiencia desde claude Opus 5.5 es qué esta atontado... Al final, vamos a ir a modelos opensource no por no pagar, si no por evitar estas bajadas de calidad :|
2
4
134
Lo bueno de Constella es que el indice ya embebido no se toca: cambias solo el encoder de las consultas. Reindexar millones de documentos cada vez que sale un modelo mejor es lo que frena a casi todos. Cuanto hace que no reindexas tu corpus?
This is a game-changer, and maybe the first of its kind. Asymmetric Dual Encoders allow you to use a normal "heavier" embedding model to encode your documents, while then giving you a "hot-swappable" model for query embeddings. You can use 400M-parameter Stella model for you documents, but on the query side choose from Zero, Nano, or Stella itself (see image for latencies). "Constella-Zero" is a learned bag of tokens that operates via look up table, pooling, and normalization. The results are shockingly good. "Constella-Nano" is a transformer model with 1/12th the parameters of Stella, and even better results Average nDCG@10 across 15 datasets: -> Zero = 0.4572 -> Nano = 0.5081 -> Stella = 0.5614 Read more about the training process, results, and getting started resources below: qdrant.tech/blog/constella-r… Credit to @DylanCouzon for the research work and @generall931 for the idea Note: Currently in "research preview"
1
255
Ayer publiqué sobre esto, sin embargo vale la pena reforzar: Citrix acaba de confirmar formalmente dos zero-days críticos (CVE-2026-88771 y CVE-2026-88772) en NetScaler ADC y Gateway, que ya estaban siendo explotados activamente antes de que saliera el parche. La decisión de...
1
2
5
440
Cada rama con su propio servidor, sus puertos y su base de datos aislada se carga el stash, reinstala y reinicia de cada cambio. Además junta los logs de todos los worktrees en una pantalla. ¿Cuántas ramas sueles tener a medias a la vez?
1
2
365
Guardar en local cada fuente que lee, en markdown, hace que la siguiente búsqueda arranque con ventaja en vez de repetir el mismo trabajo. Y al ser ficheros planos los puedes buscar tú a mano. ¿Tú guardas las fuentes de lo que investigas o las tiras al cerrar?
You can turn Claude Code into a better deep research agent than ChatGPT Deep Research It's free. It's open source. And it's called Hyperresearch The problem with every deep research tool: you get one report, then everything it read gets thrown away. Next question, it starts from zero Hyperresearch keeps EVERYTHING Every source it reads lands in a searchable vault on your computer. Plain markdown files. The next time you research anything, it checks the vault first before searching the web Your research literally compounds. Every session starts smarter than the last Here's what happens when you give it one prompt: 1. Breaks your question into pieces 2. Reads 250+ sources in a single run 3. Writes 3 drafts in parallel 4. 4 AI critics attack the report looking for weak spots 5. A patcher fixes only what the critics found. It's locked so it CAN'T rewrite the whole report 6. Every citation gets checked before it ships. Fake quotes get blocked 7. Final report, fully sourced Some details that blew me away: - 5 reprints of the same press release count as ONE source, not five - Retracted papers get flagged and blocked - Paywalled paper? It finds a legal free copy and reads the full thing - Run crashes halfway? It resumes from the exact step it died on Works in Claude Code AND Codex. 3,600+ stars on GitHub The creator says it tops the DeepResearch-Bench leaderboard in internal testing. Third party validation is still pending, but the architecture speaks for itself Setup takes 2 commands: pip install hyperresearch hyperresearch install Then type /hyperresearch followed by any question Fair warning: a full run takes 1.5 to 2.5 hours. Quick questions take 30 to 40 minutes Do yourself a favor and run it on the biggest question you have right now. Go make a coffee. Come back to a report better than anything you'd pay for github.com/jordan-gibbs/hype…
1
1
203
Un vídeo de 7 minutos que recorre la vida de una consulta deja el repo entendible antes de abrir el primer fichero. Y le sirve igual a tu propio equipo dentro de un año. ¿Qué repo te habría gustado entender así?
every single open source repo should have an explainer video like this 7min one for SQLite 1. explains what the repo does and why its useful 2. shows a high level map of the code (beautiful graphic!) 3. life of a query through the codebase 4. the core abstractions in the repo 5. a real trace of execution and what happens, with a bit on join-order query planning and yes, this was all generated by opus 5.5 w/ gemini 3.8 tts. i sound like i'm on repeat, but im in shock at just how coherent and capable this model is. everything truly is code.
1
381
Separar los papeles (uno explora, otro juzga, otro arregla) es lo que hace usable un QA automatico: el que encuentra el fallo no decide si importa. Arranca con un npx y tira de la CLI que ya tengas instalada. ¿Le dejariais abrir PRs solo, o solo reportar?
Open sourcing 𝚋𝚞𝚐𝚑𝚞𝚗𝚝𝚎𝚛𝚜 today, our internal agentic QA tool that finds and fixes bugs on autopilot. 𝚗𝚙𝚡 𝚋𝚞𝚐𝚑𝚞𝚗𝚝𝚎𝚛𝚜 𝚙𝚊𝚝𝚛𝚘𝚕 Spawns a QA team of 3 agents: - Runs forever until stopped - Detects new code changes - The "Explorer" uses your app & surfaces issues - The "Judge" triages, files detailed issues - The "Fixer" fixes and open PRs - Syncs Github, tests and verifies the fixes - Memory builds up with every run Each agent uses either your installed claude/codex/pi CLI or openrouter/vercel gateway API key. So far it supports web, electron, react native, android and ios apps. Comes with a dashboard to browse issues, watch runs live, view how the agents see your app, their memory and track token usage/cost. It's been working really great for our web, electron and expo apps, and thought other teams might find it useful! Easiest way to get started is by installing the skill 𝚗𝚙𝚡 𝚜𝚔𝚒𝚕𝚕𝚜 𝚊𝚍𝚍 𝚊𝚐𝚎𝚗𝚝–𝚕𝚊𝚋𝚜–𝚍𝚎𝚟/𝚋𝚞𝚐𝚑𝚞𝚗𝚝𝚎𝚛𝚜 Then ask your agent to "setup bughunters" in your repo. Run patrol with --once for a one time loop. Code & docs: github.com/agent-labs-dev/bu… PRs welcome!
1
326
587 ofertas comparadas: TypeScript pasa del 37% al 71% y JavaScript baja al 39%. Ya no lo piden porque lo dan por hecho: Vite, Next y Nest arrancan en TS de serie. ¿Te queda algo nuevo en JS plano?
On a JavaScript job board, fewer jobs now ask for JavaScript. I compared 587 listings from 2025 and 2026. TypeScript: 37% → 71%. JavaScript: 48% → 39%. Companies stopped writing "JavaScript". They write "TypeScript" and assume you know the rest.
2
4
299
Skill2Env convierte ficheros SKILL.md públicos en tareas que se ejecutan: 7.971 tareas en 13 dominios para entrenar agentes con RL solo por resultado. La documentación que ya escribes pasa a ser dataset de entrenamiento. ¿Aguantarían tus SKILL.md que un agente los ejecute?
NVIDIA is turning agent skills into RL environments. Skill2Env takes public SKILL.md files, turns them into executable tasks, and uses them to train agents through reinforcement learning The best part is 7,971 tasks across 13 domains. After 300 steps of outcome-only RL, Qwen3.8-27B improved from 49.4% to 54.1% on Terminal-Bench 2.1 and from 33.4% to 37.7% pass@1 on S2EBench, their hand-verified held-out benchmark Paper: github.com/NVlabs/Skill2Env/… Check this site. Where you can chat with the paper itself: academy.dair.ai/papers/reinf… Source: Dair. ai
6
182
El ahorro sale del contexto: si el agente busca y solo se trae los trozos que importan, paga menos tokens por tarea. Trae su propia skill para que el agente sepa cuándo tirar de ella. ¿Cuántos tokens se te van en leer ficheros enteros?
Introducing jevgrep - a research agent CLI powered by jev from @typesafeai that reduces your coding agent cost by 40% (verified on SWE-bench) Make sure to use the built in skill so your coding agent knows to use jg for context collection github.com/dzhng/jevgrep
3
21
1,713