# Wendelmaques — AI infrastructure consulting > AI infrastructure consulting: platforms, GPUs, model training, sensitive data, speech and integration with existing systems. Wendelmaques. Services: infrastructure review, deployment and ongoing operations. Pricing is quoted per scope. Expertise: platforms for data, training and inference; GPUs; documents and OCR; incremental synchronisation and MySQL; sensitive data and controlled agent actions; real-time audio; integration with existing systems; hybrid infrastructure, migration, observability and recovery. Contact: eu@wendelmaques.com · https://wa.me/5561991657397 The expertise pages describe implementation experience and consulting scope without identifying clients or attributing business outcomes. Capacity, availability and ongoing support are defined by the workload and proposal. ## Pages - [Consulting](https://wendelmaques.ia.br/en/): Services and contact. - [Architecture and infrastructure for AI platforms](https://wendelmaques.ia.br/en/casos/plataformas-ia/): Consulting to connect data, training and inference in an AI platform with versions, permissions, cost limits and recovery. - [AI architecture for sensitive data](https://wendelmaques.ia.br/en/casos/ia-dados-sensiveis/): Consulting for AI with sensitive data: organisation-scoped access, protected media, audit records and source checks before human review. - [AI integration with critical systems and telephony](https://wendelmaques.ia.br/en/casos/integracao-sistemas-criticos/): Consulting to integrate AI with existing systems: C++/Qt, MySQL, APIs, events and SIP telephony, with staged deployment and recovery. - [LLMs and speech on one GPU: organising operations](https://wendelmaques.ia.br/en/casos/orquestracao-gpu/): Consulting to assess GPU sharing, coordinate inference services and define capacity and waiting-time limits. - [Private vLLM inference for your product](https://wendelmaques.ia.br/en/casos/inferencia-vllm/): Consulting to deploy private inference, connect the API to your product and organise capacity, access and operations. - [Whisper transcription and speech service operations](https://wendelmaques.ia.br/en/casos/transcricao-whisper/): Consulting for real-time audio, WebRTC, Opus and Whisper transcription: capture, queues, media protection and processing recovery. - [AI in production: what you need beyond the model](https://wendelmaques.ia.br/en/artigos/ia-em-producao/): A deployment guide for companies: real workloads, APIs, queues, access, monitoring, versions and recovery. - [How to size GPUs and AI inference costs](https://wendelmaques.ia.br/en/artigos/capacidade-gpu-ia/): Context, concurrency, queueing and latency: measurement criteria for planning an AI product's infrastructure and costs. - [How to prepare AI for sensitive data](https://wendelmaques.ia.br/en/artigos/ia-com-dados-sensiveis/): Define access, data flows, storage, logs and human review before connecting AI to restricted documents and conversations. - [How to connect AI to the systems your company already uses](https://wendelmaques.ia.br/en/artigos/integrar-ia-sistemas-existentes/): Plan AI integration with databases, applications and telephony: contracts, action confirmation, queues, recovery and staged deployment. - [Data quality before training an AI model](https://wendelmaques.ia.br/en/artigos/qualidade-dados-treinamento-ia/): Versions, sources, duplicates and separation between training and evaluation: how to prepare data for a training run you can verify. - [Distributed AI training: budgets, ownership and recovery](https://wendelmaques.ia.br/en/artigos/treinamento-distribuido-modelos-ia/): How to control long training jobs when workers lose connection, GPUs become unavailable and spending needs a limit. - [How to verify and release a new model version](https://wendelmaques.ia.br/en/artigos/promocao-versoes-modelos-ia/): Separate training, file checks, evaluation and release to control which versions reach inference. - [Documents for AI: extraction, OCR and verifiable quality](https://wendelmaques.ia.br/en/artigos/documentos-ocr-ia/): How to prepare PDFs and documents for search and AI without hiding missing pages, lost structure or text recognition failures. - [Data pipelines for AI: sync changes with recovery](https://wendelmaques.ia.br/en/artigos/pipelines-dados-incrementais/): Versions, deletions, bounded batches and destination confirmation: decisions for updating AI data without reloading the whole collection. - [MySQL in AI products: reduce work before enlarging the server](https://wendelmaques.ia.br/en/artigos/mysql-custo-operacao-ia/): Frequent queries, indexes and small writes can dominate cost. Measure daily work and address the causes of reads and writes. - [AI queues: retries, concurrency and external effects](https://wendelmaques.ia.br/en/artigos/filas-tarefas-ia/): How to recover AI tasks without letting old executions confirm results or a lost response cause duplicate actions. - [AI observability: find where operations stopped](https://wendelmaques.ia.br/en/artigos/observabilidade-infraestrutura-ia/): Connect errors, queues, versions and physical resources to delivered results. An available service can still fail to produce useful work. - [Release AI services with a planned recovery path](https://wendelmaques.ia.br/en/artigos/publicacao-recuperacao-servicos-ia/): Check the new service, compatibility and loaded model before switching traffic. Plan recovery beyond the container image. - [Backup and recovery for an AI platform](https://wendelmaques.ia.br/en/artigos/backup-recuperacao-plataformas-ia/): Databases, documents, models and queue state need a recovery procedure. Replication and a backup file do not prove recovery. - [Real-time audio for AI: capture before processing](https://wendelmaques.ia.br/en/artigos/audio-tempo-real-ia/): Transport, recording and transcription have different failures. Separate these stages to allow recovery and show the actual state of each segment. - [Hybrid AI infrastructure: where to place data and processing](https://wendelmaques.ia.br/en/artigos/infraestrutura-hibrida-ia/): Distribute processing without making every operation a network crossing. Database location, GPUs and access limits guide the design. - [Migrate and retire a stack while keeping operational control](https://wendelmaques.ia.br/en/artigos/migracao-infraestrutura-ia/): Verify the destination, preserve data and assign a single owner to each resource before retiring old infrastructure. - [AI agents: separate interpretation from permission to act](https://wendelmaques.ia.br/en/artigos/acoes-controladas-agentes-ia/): How to connect AI to business systems with bounded actions, context-specific access and result checks outside the model. - [Personal Agent Protocol: what you can prepare now](https://wendelmaques.ia.br/en/artigos/personal-agent-protocol/): The PAP specification is not out yet. PACT has public documentation. Prepare discovery, authorisation and billing without claiming an integration that does not exist. - [Agent Access Discovery: specification 0.1](https://wendelmaques.ia.br/en/artigos/especificacao-acesso-agentes/): Independent draft for interfaces, documentation and prices. Discovery does not grant access. This implementation excludes PAP and PACT. - [Wendelmaques: AI platforms, integration and infrastructure](https://wendelmaques.ia.br/en/experiencia/): Over 12 years in engineering. AI platforms, sensitive data, training, real-time audio and integration with existing systems. - [AI infrastructure articles](https://wendelmaques.ia.br/en/artigos/): Engineering decisions for data, models, speech and operations. Articles by Wendelmaques based on implementation experience. - [Frequently asked questions](https://wendelmaques.ia.br/en/faq/): Answers about the AI infrastructure consulting: services, pricing, steps and contact. - [Wendelmaques.com account](https://wendelmaques.ia.br/en/conta/): Access to WRP applications. Understand Google sign-in and integration permissions before you authorise access. - [Account privacy](https://wendelmaques.ia.br/en/conta/privacidade/): The data your account uses, why it is needed and how to exercise your rights. - [Account terms](https://wendelmaques.ia.br/en/conta/termos/): Conditions for access to WRP applications and permission for integrations. - [Briefs: news and security advisories, with what to do](https://wendelmaques.ia.br/en/pautas/): Each brief summarizes a technology news item or a security advisory and explains what a company can do in response. - [Unpatched vulnerabilities in Cyrus SASL and no new release since 2022](https://wendelmaques.ia.br/en/pautas/vulnerabilidades-sem-correcao-no-cyrus-sasl-d1774f/): Alan Coopersmith posted on the oss-sec list that Cyrus SASL has publicly disclosed security problems with no response from maintainers, and the project has not shipped a release since 2022. - [CVE-2026-102916: rep prefix emulation flaw in illumos bhyve](https://wendelmaques.ia.br/en/pautas/cve-2026-102916-falha-na-emulacao-2a06b0/): illumos disclosed CVE-2026-102916, which affects rep prefix instruction emulation in bhyve, the VMM hypervisor. The flaw mishandles flags and requires patching in downstream distributions. - [Python 3.15: sentinel, frozendict, UTF-8 by default and experimental JIT](https://wendelmaques.ia.br/en/pautas/python-3-15-sentinel-frozendict-utf-9dd3e2/): LWN reports the release of Python 3.15, with the built-in sentinel and frozendict types, UTF-8 as the default encoding and improvements to the experimental JIT compiler. See what changes for teams maintaining Python systems. - [Control-flow integrity checks for the Linux kernel added to GCC](https://wendelmaques.ia.br/en/pautas/verificacao-de-integridade-do-fluxo-de-37b1de/): Kees Cook presented changes to GCC at the 2026 GNU Tools Cauldron that add forward-edge CFI support for the kernel, hindering exploits that divert execution from its intended paths. - [Data breach at iRhythm affects hundreds of thousands of people](https://wendelmaques.ia.br/en/pautas/vazamento-de-dados-na-irhythm-afeta-4f638d/): iRhythm, maker of wearable cardiac sensors, has begun notifying states about a data breach that occurred in the summer. The Record reports impact on hundreds of thousands of people, showing how late detection widens the risk for companies that handle health data. - [Artificial Analysis: HeyGen Voice tops controlled voice TTS ranking](https://wendelmaques.ia.br/en/pautas/artificial-analysis-heygen-voice-lidera-ranking-15c25d/): Artificial Analysis ranked HeyGen Voice first in the Controlled Voice TTS Arena, with an Elo of 1,201 across 8 cloned voices in US and UK English, ahead of competing commercial voice models. - [HERMES: component-level harness with a resident LLM per code module](https://wendelmaques.ia.br/en/pautas/hermes-harness-em-nivel-de-componente-ed1773/): An arXiv paper proposes HERMES, where each repository component has a resident LLM that knows its code, with reported gains in migration and on four software engineering benchmarks. - [Qwen3-TTS 1.7B: open-source voice cloning with emotion and accent controls](https://wendelmaques.ia.br/en/pautas/qwen3-tts-1-7b-clonagem-de-3d31d6/): The post presents an open-source 1.7B text-to-speech model for voice cloning, with controls for emotion, age, tone and accent, plus multilingual support. For companies, a self-hosted evaluation means checking license, languages and controls before any use. - [Whistle: 16.9 MB speech-to-text model built for on-device use](https://wendelmaques.ia.br/en/pautas/whistle-modelo-de-speech-to-text-869c9e/): Cactus-Compute published Whistle, a 16.9 MB speech-to-text model for on-device use. The author claims it rivals Whisper base. For companies handling sensitive audio, the announcement opens a case for local transcription. - [AI training deal involving Spirit Airlines data draws lawmakers' warning](https://wendelmaques.ia.br/en/pautas/acordo-de-treinamento-de-ia-com-3ff6fa/): Lawmakers warn that a proposed sale of Spirit Airlines assets would include about 100 million emails, 500 million Microsoft Teams messages and employment and financial records, with risk of use in AI training. - [Ubuntu advisory USN-8910-1: libxml2 flaws with risk of denial of service, XXE and SSRF](https://wendelmaques.ia.br/en/pautas/aviso-ubuntu-usn-8910-1-falhas-d14fe0/): The Ubuntu security advisory USN-8910-1 describes six vulnerabilities in libxml2, with risks of denial of service, code execution, XML external entity injection and server-side request forgery. One of them affects only Ubuntu 26.04 LTS. - [Go project advisory: HTTP/2 server memory exhaustion fixed in golang.org/x/net v0.60.0](https://wendelmaques.ia.br/en/pautas/aviso-do-projeto-go-exaustao-de-1d4cfa/): The Go project published a security advisory for golang.org/x/net v0.60.0, which fixes several vulnerabilities, including memory exhaustion in HTTP/2 servers related to Trailer headers. - [Go 1.27.2 and 1.26.9 fix 15 security issues](https://wendelmaques.ia.br/en/pautas/go-1-27-2-e-1-cf337d/): The Go project released versions 1.27.2 and 1.26.9 with 15 security fixes. Here is what this requires from teams running Go services. - [International coalition disables tools of Integrity Tech linked to Flax Typhoon](https://wendelmaques.ia.br/en/pautas/coalizao-internacional-desativa-ferramentas-da-integrity-239320/): The United States and other countries took down digital tools and infrastructure of Integrity Tech, a Beijing-based company, used for widespread vulnerability scanning and, in some cases, intrusions in the Flax Typhoon campaign. - [Text extracted from malware samples can manipulate AI-assisted analysis](https://wendelmaques.ia.br/en/pautas/textos-extraidos-de-amostras-de-malware-8755e6/): Cisco Talos highlights malware techniques that try to manipulate AI-assisted analysis and recommends treating text extracted from samples as evidence, not instructions. Companies using models for triage need to separate data from commands. - [USN-8909-1: denial-of-service flaw in the libde265 H.265 decoder](https://wendelmaques.ia.br/en/pautas/aviso-usn-8909-1-falha-de-946fb8/): Ubuntu Security published an advisory on a validation flaw in certain malformed H.265 bitstreams in libde265, leading to a NULL pointer dereference that can crash the decoder and cause denial of service. For companies processing HEVC video, the risk is interruption of media services. - [How the Rust project organizes, tracks and reviews its goals](https://wendelmaques.ia.br/en/pautas/como-o-projeto-rust-organiza-acompanha-c53e54/): Tomáš Šedovič, program manager at the Rust Foundation, presented at Kangrejos 2026 the goals process of the Rust project, aimed at Rust for Linux developers. A practical view of governance for long-running technical initiatives. - [AI agent forensics: where conversation history lives in OpenCode and Hermes](https://wendelmaques.ia.br/en/pautas/forense-de-agentes-de-ia-onde-e883b8/): The SANS ISC post examines how to search for OpenCode and Hermes conversation history as evidence in investigations of AI use, within the FOR577 course update. - [Dispute in the Linux kernel over swap cache slots and persistent swap file space](https://wendelmaques.ia.br/en/pautas/disputa-no-kernel-linux-sobre-slots-675a68/): Two competing proposals address the direct link between swap cache slots and space in persistent swap files, with no consensus yet for merging into the mainline kernel. - [DeLM: decentralized code agents with shared context, without a central coordinator](https://wendelmaques.ia.br/en/pautas/delm-agentes-de-codigo-descentralizados-com-e1bdac/): Stanford's article presents DeLM, which replaces a central coordinator with shared context and a task queue. According to the author, it is up to 2.49× faster and more accurate than the Claude Code and Codex baselines. - [LayerRoPE: encoding depth in normalization weights for very deep Transformers](https://wendelmaques.ia.br/en/pautas/layerrope-codificar-a-profundidade-nos-pesos-28aed1/): The paper proposes LayerRoPE, which encodes depth in normalization weights to contain the growth of hidden-state norm across layers, reporting the loss of a 1.3B Pre-Norm model with 3.4x less compute and stable scaling up to 512 layers. - [RoboJEPA: how scaling JEPA world models affects robotic planning](https://wendelmaques.ia.br/en/pautas/robojepa-como-a-escala-de-modelos-e641b8/): Meta presented RoboJEPA, a latent world model based on JEPA, scaled from 22M to 8B parameters and trained on data from 12 robotic embodiments. The paper tests whether latent prediction error predicts later planning performance. - [Samsung opens LittleBit, a method that compresses 13B LLMs to under 1 GB](https://wendelmaques.ia.br/en/pautas/samsung-abre-codigo-do-littlebit-metodo-e16e43/): According to the post, Samsung's LittleBit uses latent factorization and sub-1-bit levels to shrink the weights of 13B LLMs to under 1 GB. Worth evaluating for local deployment. - [CVE-2026-97146: YuniKorn admission control bypass](https://wendelmaques.ia.br/en/pautas/apache-yunikorn-falha-contorna-controle-de-3f9337/): An OSS-Sec advisory published on October 7, 2026 describes a way to bypass annotation checks in Apache YuniKorn 1.9.0 and earlier. The reported severity is medium, with a CVSS 4.0 score of 4.8. - [Apache YuniKorn: admission check bypass](https://wendelmaques.ia.br/en/pautas/cve-2026-92393-falha-no-yunikorn-04b28b/): An advisory published on October 7, 2026, reports that Apache YuniKorn 1.9.0 and earlier do not check user labels and annotations for workload UPDATE operations, allowing admission controls to be bypassed. The reported severity is low (CVSS 4.0: 2.0). - [Multilingual speech-to-text benchmark for local models in transcribe.cpp on a notebook](https://wendelmaques.ia.br/en/pautas/benchmark-multilingue-de-speech-to-text-344d73/): Site comparing accuracy and speed of speech-to-text models supported by transcribe.cpp, measured per language on FLEURS, with performance data from a notebook CPU/GPU (4750U). - [Voice clones of non-native English speakers rated more human: what it means for companies](https://wendelmaques.ia.br/en/pautas/clones-de-voz-de-falantes-nao-a612f7/): A NeurIPS 2026 paper cloned the voices of 86 non-native English speakers using three voice-synthesis systems. In 4,000 human ratings, the clones were judged warmer, more authoritative and more human, which raises trust risks for voice channels. - [H-JEPA: hierarchical world models for long-horizon visual planning](https://wendelmaques.ia.br/en/pautas/h-jepa-modelos-de-mundo-hierarquicos-b23a6e/): The article, guided by Yann LeCun, proposes H-JEPA, with JEPA world models stacked across multiple time scales. The author reports a rise from 18% to 73% on Visual AntMaze with this hierarchical approach. - [LiFT: looped transformers for image generation with flow matching and adjustable inference cost](https://wendelmaques.ia.br/en/pautas/lift-transformers-em-loop-para-gerar-b3010f/): The paper presents LiFT, an image generator with flow matching that reuses a shared transformer core across loops. According to the author, a model trained with 2 loops improves when run with up to 16 loops at inference, without retraining. - [Memory-Efficient Expert Routing for Distributed MoE Training](https://wendelmaques.ia.br/en/pautas/roteamento-de-experts-com-eficiencia-de-6b779e/): The arXiv paper proposes methods to reduce the memory peak in distributed Mixture-of-Experts training, where expert dispatch dominates memory use. - [Single-voice 9MB TTS model distilled from Kokoro-82M for local voice output](https://wendelmaques.ia.br/en/pautas/tts-de-voz-unica-com-9mb-f087a0/): A post describes a single-voice, single-language text-to-speech model of about 9MB of weights, distilled from Kokoro-82M and aimed at local voice output on resource-constrained devices. - [Ubuntu warns of token exposure in WSL](https://wendelmaques.ia.br/en/pautas/usn-8892-1-token-exposto-no-82e512/): An Ubuntu Security advisory published on October 6, 2026 reports that the Ubuntu Pro for WSL attachment token could appear in command-line arguments during subscription activation. - [Ubuntu fixes libsoup flaws with exposure and denial-of-service risks](https://wendelmaques.ia.br/en/pautas/usn-8890-1-falhas-no-libsoup-03f746/): Ubuntu Security notice USN-8890-1, published on October 6, 2026, describes libsoup flaws that could enable denial of service, information disclosure, security-control bypass, and code execution. Fixes are available for several Ubuntu versions. - [CVE-2026-104380: routing flaw in Punk for Perl](https://wendelmaques.ia.br/en/pautas/cve-2026-104380-falha-no-roteamento-2a61d1/): Advisory published on the oss-sec mailing list on October 5, 2026: Punk for Perl versions 0.48 through releases before 0.55 may route Extended CONNECT to GET routes without validating Origin. - [On-policy distillation combines specialist models](https://wendelmaques.ia.br/en/pautas/destilacao-on-policy-combina-modelos-especialistas-a2c112/): A post published on October 6, 2026 describes MOPD, a method that trains a student model on its own rollouts with token-level supervision from domain specialists. - [JEPA-Anything explores predictive learning across seven domains](https://wendelmaques.ia.br/en/pautas/jepa-anything-explora-aprendizado-preditivo-em-7a3e0b/): The post, published on October 6, 2026, presents JEPA-Anything as a domain-agnostic world-modeling framework based on orthogonal predictive factorization. - [LiFT expands iterative inference with a shared DiT core](https://wendelmaques.ia.br/en/pautas/lift-amplia-a-inferencia-iterativa-com-e5484d/): Published on October 6, 2026, the research reports that LiFT outperforms a dense DiT baseline on ImageNet 256×256 with about 60% fewer parameters. - [Quantizing 10M EmbeddingGemma vectors to 1 bit with MRL cuts RAM from 30GB to 0.4GB](https://wendelmaques.ia.br/en/pautas/quantizacao-de-10m-de-vetores-embeddinggemma-76a60a/): The author's post reports compressing 10 million EmbeddingGemma embeddings to 1 bit, with MRL reduction to 256 dimensions, dropping RAM from about 30GB to 0.4GB at a stated quality loss of about 5%. - [Multimodal retrieval with embeddings in a shared space](https://wendelmaques.ia.br/en/pautas/retrieval-multimodal-com-embeddings-em-um-29008d/): In a post published on October 6, 2026, Google says EmbeddingGemma 2 represents text, code, images, video, and audio in a shared vector space and uses the Apache 2.0 license. - [AI agents test the limits of web platforms](https://wendelmaques.ia.br/en/pautas/wikimedia-alerta-sobre-agentes-de-ia-10c1f8/): On October 5, 2026, Wikimedia reported attempts by AI agents to edit pages and compromise a note-taking tool, as well as resource consumption on web platforms. - [libexpat 2.9.0 fixes two vulnerabilities](https://wendelmaques.ia.br/en/pautas/libexpat-2-9-0-corrige-duas-ccbc8d/): A post on the oss-security list says libexpat 2.9.0 fixes two vulnerabilities, including an integer overflow on 32-bit platforms. - [Early Patches Help Cloud Providers Prepare Mitigations](https://wendelmaques.ia.br/en/pautas/correcoes-antecipadas-e-mitigacao-na-nuvem-581709/): On October 5, 2026, a reply posted to oss-sec argued that giving cloud providers patches before public advisories helps schedule mitigations and identify gaps in patches. - [Pwn2Own Ireland 2026: three days of testing](https://wendelmaques.ia.br/en/pautas/pwn2own-ireland-2026-tres-dias-de-4d1478/): Published on October 5, 2026, the schedule announces more than 60 entries across three days, with contests targeting devices and systems in several categories, including AI infrastructure. - [New NetScaler alert calls for an exposure review](https://wendelmaques.ia.br/en/pautas/eua-e-australia-alertam-sobre-novo-6a994d/): The Record reports that the United States and Australia warned of a newly observed issue in customer-managed NetScaler deployments, distinct from the vulnerabilities disclosed the previous week. - [Automated patch review: lessons from Sashiko](https://wendelmaques.ia.br/en/pautas/sashiko-automatiza-revisoes-de-patches-do-6e3518/): LWN's article presents Sashiko, a system that uses a language model to review kernel patches automatically and help address a shortage of reviewers. - [Ubuntu warns of authorization flaws in Aodh and Watcher](https://wendelmaques.ia.br/en/pautas/usn-8870-1-autorizacao-no-aodh-c51d77/): On October 5, 2026, Ubuntu Security published an advisory about flaws in OpenStack Aodh and Watcher that could expose alarm metadata and allow unauthorized triggers. - [Ubuntu fixes Raspberry Pi Linux kernel vulnerabilities](https://wendelmaques.ia.br/en/pautas/usn-8871-1-falhas-no-kernel-c67988/): Ubuntu Security advisory USN-8871-1, published on October 5, 2026, covers Linux kernel vulnerabilities for Raspberry Pi, including a risk on Arm processors and issues across several subsystems. - [Linux kernel flaws for Azure call for update reviews](https://wendelmaques.ia.br/en/pautas/falhas-no-kernel-linux-para-azure-7182f1/): Security advisory USN-8851-3 reports fixes in the NFS server, IPv6 networking and Netfilter for the Linux kernel for Azure, addressing three CVEs. - [October 5 report highlights data, cloud and AI risks](https://wendelmaques.ia.br/en/pautas/relatorio-de-5-de-outubro-destaca-7fe7d2/): Check Point Research’s report, published on October 5, 2026, covers data breaches, ransomware, exploited vulnerabilities and AI-agent activity. The cases point to practical steps for protecting data, backups and cloud environments. - [TTY logs help correlate attackers’ commands](https://wendelmaques.ia.br/en/pautas/registros-tty-revelam-comandos-apos-invasoes-9308a5/): In an article published by SANS ISC on October 5, 2026, an experiment sends command logs to the DShield SIEM daily after attackers or bots gain access to a sensor. - [Attention sinks and rescaling in transformers](https://wendelmaques.ia.br/en/pautas/attention-sinks-e-reescala-em-transformers-4a2fc6/): Published on October 5, 2026, the paper examines attention sinks and residual sinks in large language models and proposes how these outliers, softmax attention, and RMSNorm rescale other components. - [Operational tags help prioritize SIEM detection rules](https://wendelmaques.ia.br/en/pautas/como-classificar-regras-de-deteccao-por-3a62a5/): In an article published on October 5, 2026, Elastic Security Labs describes monthly tags for assessing noise, performance, threat, and recommendation across preconfigured SIEM rules. - [Test training scale before expanding GPU capacity](https://wendelmaques.ia.br/en/pautas/como-testar-a-escala-de-treinamento-00eb13/): In a report published on October 5, 2026, Aleph Alpha describes scaling a 30B-A3B MoE model from 16 to 512 B200 GPUs, reporting near-linear scalability and 35.3% MFU. - [Dust explores gradient-free transformer training](https://wendelmaques.ia.br/en/pautas/dust-explora-treinamento-de-transformers-sem-2555c8/): Published on October 5, 2026, the paper presents Dust, a zeroth-order method using activation perturbations, and reports scaling experiments and comparisons with backpropagation. - [GPU indexes expand PyTorch and Python support](https://wendelmaques.ia.br/en/pautas/indices-de-gpu-ampliam-suporte-a-9d4dfe/): An Astral post published on October 5, 2026, reports support for PyTorch 2.13 and 2.14 and Python 3.15, with GPU wheels for packages including FlashAttention and DeepSpeed. - [Shared memory reduces memory use in recurrent Transformers](https://wendelmaques.ia.br/en/pautas/memoria-compartilhada-reduz-o-uso-de-4b1438/): An article published on October 5, 2026 reports a 76–79% reduction in context memory and improved quality in models with 150 million to 1 billion parameters. - [Native Hybrid Attention and the long-context trade-off](https://wendelmaques.ia.br/en/pautas/native-hybrid-attention-e-o-equilibrio-8a32d4/): An article published on October 5, 2026 presents Native Hybrid Attention, an approach aimed at balancing speed in long contexts with detail retention. Companies can assess the proposal using benchmarks tied to their tasks. - [SwiLA proposes switching between linear attention maps](https://wendelmaques.ia.br/en/pautas/swila-propoe-alternar-entre-mapas-lineares-814659/): A paper presented as work at COLM 2026 and dated October 5, 2026 describes an approach that switches among multiple linear maps and highlights the trade-off between expressiveness and memory. - [Symmetric memory may optimize GPU communication](https://wendelmaques.ia.br/en/pautas/symmetric-memory-pode-otimizar-comunicacao-em-8c9d61/): A Machine Learning Engineering Open Book post, published on October 5, 2026, discusses symmetric memory in NCCL and PyTorch and its potential for small and medium payloads and overlapping communication with computation. - [Vulnerability disclosure for cloud providers](https://wendelmaques.ia.br/en/pautas/divulgacao-de-vulnerabilidades-para-provedores-de-f25c78/): In a post dated October 4, 2026, oss-sec proposes a disclosure list for cloud and VPS hosting providers that do not participate in Linux distribution lists. - [Apache Thrift 0.25.0 fixes 61 vulnerabilities](https://wendelmaques.ia.br/en/pautas/apache-thrift-0-25-0-corrige-3a3d38/): The oss-sec advisory says Apache Thrift 0.25.0 fixes 61 vulnerabilities in earlier versions and recommends upgrading. - [CISA adds NetScaler flaw to KEV catalog](https://wendelmaques.ia.br/en/pautas/cisa-inclui-falha-do-netscaler-no-eb8d56/): On October 4, 2026, CISA added CVE-2026-88779, a Citrix NetScaler flaw, to its KEV catalog following evidence of active exploitation. - [Speaker diarization: evaluating label correction](https://wendelmaques.ia.br/en/pautas/diarizacao-de-fala-como-avaliar-a-580379/): A Google article published on October 4, 2026 describes a 4-billion-parameter model fine-tuned to correct speaker labels and reduce word-level diarization errors across four benchmarks. - [BF16 gradients rise late in training](https://wendelmaques.ia.br/en/pautas/gradientes-bf16-crescem-no-fim-do-25485c/): A paper reports a 1,000-fold increase in gradient norm after 25 billion tokens in a transformer trained with FlashAttention-3 in BF16. - [Projection Sampling adapts demonstrations for SFT](https://wendelmaques.ia.br/en/pautas/projection-sampling-adapta-demonstracoes-para-sft-f75426/): A paper published on October 4, 2026 presents a way to rewrite expert demonstrations as correct task trajectories that are more likely under the target model, then use them in standard supervised fine-tuning. - [Router-R1 explores reinforcement learning to route LLMs](https://wendelmaques.ia.br/en/pautas/router-r1-e-o-roteamento-de-e80fcc/): Published on October 18, 2025, the text presents Router-R1, which alternates between reasoning and calls to multiple models and reports evaluations on seven question-answering benchmarks. - [Audio codec for Speech LLMs: parallel tokens and streaming](https://wendelmaques.ia.br/en/pautas/codec-de-audio-para-speech-llms-49231d/): In a post published on October 17, 2025, Meituan presented an open-source audio codec optimized for Speech LLMs, with parallel semantic and acoustic tokens at 16.7 Hz and low-latency streaming decoding. - [A Python RLM implementation for long contexts](https://wendelmaques.ia.br/en/pautas/rlm-open-source-analisa-contextos-longos-aeeb06/): Published on October 17, 2025, this repository presents an open-source Recursive Language Models implementation that keeps long inputs in a Python environment for analysis. - [AGI: definitions and benchmarks merit critical analysis](https://wendelmaques.ia.br/en/pautas/agi-definicoes-e-limites-dos-benchmarks-214ef1/): A publication dated October 16, 2025, discusses a paper defining AGI as the replication of human cognition. Its author welcomes discussion of varying capabilities but questions the benchmarks and the narrow view of AI. - [Reflective prompt optimization with DSPy GEPA](https://wendelmaques.ia.br/en/pautas/dspy-gepa-otimizacao-reflexiva-de-prompts-864faa/): A cookbook published on October 16, 2025 presents reflective prompt optimization with DSPy GEPA and reports an 11% accuracy gain on NuminaMath-1.5 for less than US$0.50 in total. - [Parallel algorithms and memory hierarchy](https://wendelmaques.ia.br/en/pautas/algoritmos-paralelos-e-hierarquia-de-memoria-0e5238/): Guy Blelloch’s book covers parallel algorithms and memory hierarchy, topics relevant to assessing the performance of systems that process tasks in parallel. - [RLMs: an approach to very long prompts](https://wendelmaques.ia.br/en/pautas/rlms-prompts-longos-tratados-por-meio-0b0fc0/): Published on October 15, 2025, the publication describes language models that interact recursively with long prompts through a REPL and reports evaluations on the OOLONG and BrowseComp-Plus benchmarks. - [Radxa Orion O6N combines 12 cores with an AI accelerator](https://wendelmaques.ia.br/en/pautas/radxa-orion-o6n-placa-nano-itx-586d61/): Published on October 14, 2025, the article describes a Nano-ITX single-board computer with a CIX P1 SoC, up to 64 GB of LPDDR5, and a 30/45 TOPS AI accelerator. - [SkyRL tx v0.0.2 enables multi-LoRA training](https://wendelmaques.ia.br/en/pautas/skyrl-tx-v0-0-2-e-68a3b4/): On October 14, 2025, version 0.0.2 of the open SkyRL tx backend added support for training multi-LoRA models through the Tinker API. - [ACE: structured context that evolves with feedback](https://wendelmaques.ia.br/en/pautas/ace-acumula-contexto-com-feedback-de-4b6125/): A post dated October 13, 2025 describes Agentic Context Engineering, an approach that incrementally accumulates and refines structured context using performance feedback as an alternative to iterative prompt rewriting. - [Audio reasoning benchmark reports 92%](https://wendelmaques.ia.br/en/pautas/audio-com-raciocinio-qualidade-e-latencia-3b4f2c/): An Artificial Analysis report records a 92% score on Big Bench Audio for an audio model and compares time to first token with and without reasoning. - [InfLLM-V2 switches between dense and sparse attention](https://wendelmaques.ia.br/en/pautas/infllm-v2-alterna-atencao-densa-e-2b1832/): Published on October 13, 2025, the paper presents an attention system that switches between dense and sparse modes for short and long sequences. - [Study compares attention in 13 LLMs, with 17% performing best](https://wendelmaques.ia.br/en/pautas/estudo-compara-attention-em-13-llms-ce50f3/): Published on October 12, 2025, the study reports experiments training 13 LLMs with different proportions of conventional attention and DeltaNet linear attention. The configuration with 17% attention—two of 12 layers—achieved the best result. - [LLM post-training: methods and evaluation](https://wendelmaques.ia.br/en/pautas/guia-de-post-training-de-llms-de11c7/): A guide published on October 12, 2025, on adapting LLMs: from next-token prediction to instruction following, including training methods and evaluation. - [High Bandwidth Flash: what to assess before 2026](https://wendelmaques.ia.br/en/pautas/hbf-mais-capacidade-de-memoria-para-a31f68/): In an article published on October 12, 2025, SanDisk proposes NAND-based memory with 8 to 16 times the capacity of HBM. First samples are expected in the second half of 2026. Portuguese pages are at the same paths without /en. Sitemap: https://wendelmaques.ia.br/sitemap.xml