IBM at ICML 2026

  • Seoul, South Korea
This event has ended.

About

IBM is proud to be sponsoring the 43rd International Conference on Machine Learning (ICML).

The ICML is the premier gathering of professionals dedicated to the advancement of the branch of artificial intelligence known as machine learning. ICML is globally renowned for presenting and publishing cutting-edge research on all aspects of machine learning used in closely related areas like artificial intelligence, statistics and data science, as well as important application areas such as machine vision, computational biology, speech recognition, and robotics. ICML is one of the fastest growing artificial intelligence conferences in the world. Participants at ICML span a wide range of backgrounds, from academic and industrial researchers, to entrepreneurs and engineers, to graduate students and postdocs.


Booth Information

Visit us at the IBM booth in the exhibit hall to talk to our researchers and see demos of our work.

IBM Booth Talks & Demos:

  • Composable AI: Build AI More Like Software (Talk)
  • The Modular Engine of Enterprise AI: Precision, efficiency, and trust with the IBM Granite 4.1 Model Family (Talk)
  • Granite Libraries + Mellea
for Modular LLM Programming (Demo)
  • Granite Speech + Mellea for Real-Time Voice Interaction (Demo)
  • Granite Vision + Docling for Document Intelligence (Demo)
  • IBM Project Bob (Demo)

Agenda

  • Description:

    vLLM-Hook is a modular plug-in library for vLLM that lets developers and researchers inspect, analyze, and intervene on internal model states during inference. The talk will present the core design of vLLM-Hook, including its configuration-driven hook interface, support for passive programming and active programming, and compatibility with practical deployment workflows. We will show how the system exposes internal signals such as attentions, attention heads, and activations, and how these signals can be used for real-time monitoring and controlled intervention without requiring model retraining. The session will highlight three concrete use cases from the project: prompt-injection detection through in-model monitoring, retrieval enhancement through selective retrieval and reranking signals, and activation steering for controlled generation. The goal of the talk is to give practitioners a clear view of how model-internal programming can become a practical capability in modern LLM serving stacks built on vLLM.

    https://icml.cc/virtual/2026/75729 https://github.com/IBM/vLLM-Hook

    Speakers:
    PC
    Chief Scientist, RPI-IBM AI Research Collaboration; Research Staff Member - Adversarial Machine Learning
    IBM
    KN
    Kenney Ng
    Principal Research Scientist, IBM Research | Science Program Manager, MIT-IBM Computing Research Lab
    UP
    Unknown Person
  • Description:

    We are demonstrating a live voice assistant, built on open IBM Granite 4.1 models, that lets ICML attendees watch a language model check its own work in real time, turn by turn, during a natural spoken conversation. Attendees walk up, speak to the assistant, and watch a panel beside it light up as the system generates several candidate responses in parallel and scores each one against a set of plain-English requirements: how the answer should sound, how long it should be, what it shouldn't say. Passing requirements turn green, failures turn red, and the first candidate that satisfies all of them is spoken back. Failures are shown, not hidden. Attendees can edit the requirements on the fly and hear the assistant's behavior change mid-conversation. The demonstration is hands-on and built for a research audience working on validated and controllable generation. Every piece of it (models, orchestration, frontend) is Apache-2.0 and runs on a single laptop with no external API, so any attendee can reproduce it after the session.

    https://icml.cc/virtual/2026/75718 https://github.com/generative-computing/mellea-demos/tree/main/2026-granite-speech

    Speakers:
    KN
    Kenney Ng
    Principal Research Scientist, IBM Research | Science Program Manager, MIT-IBM Computing Research Lab
    HL
    Heiko Ludwig
    Principal Research Staff Member, Senior Manager AI Foundations Engineering
    UP
    Unknown Person
  • Description:

    Every LLM application eventually runs into the same wall: the model generates plausible-sounding output that is wrong, off-format, or unsafe — and there is nothing between generation and delivery to catch it. Prompting the model harder helps sometimes. However, it is not reliable.

    This workshop teaches a systematic approach to the problem using two open-source IBM tools: Mellea, a Python library for structured LLM generation, and Granite Libraries, a collection of lightweight LoRA adapters that score generated output against developer-defined requirements. Together they implement an Instruct-Validate-Repair loop — generate a response, measure it against your requirements, and select or retry before it reaches the user.

    Participants start with a plain chatbot that hallucinates citations, ignores formatting rules, and produces uncontrolled output. By the end of the session, the same chatbot validates every response against a set of requirements defined in plain English, generates multiple candidates in parallel, and automatically selects the best one — all running locally, all on open-source models.

    No cloud accounts, no audio hardware, no frontend build. A working environment takes under five minutes to set up.

    What you will build: a command-line chat application that grows module by module — from a bare Mellea generation call, to single-requirement scoring, to a parallel Best-of-N validation loop with multiple Granite Libraries adapters firing simultaneously.

    What you will leave with: a mental model of how to enforce output quality programmatically, hands-on experience writing and tuning natural-language requirements, and a local codebase you can adapt to your own domain.

    Technologies covered: Mellea, Granite Libraries (activated LoRA adapters), IBM Granite 4.0, Python, OpenAI-compatible inference backends (LM Studio, Ollama, vLLM).

    All tools and models used are Apache 2.0 licensed and available on HuggingFace.

    https://icml.cc/virtual/2026/75711 https://mellea.ai/

    Speakers:
    JL
    Jake LoRocco
    IBM
    KN
    Kenney Ng
    Principal Research Scientist, IBM Research | Science Program Manager, MIT-IBM Computing Research Lab
    HL
    Heiko Ludwig
    Principal Research Staff Member, Senior Manager AI Foundations Engineering
    UP
    Unknown Person

Careers @ IBM

Visit us at the IBM Booth to meet with IBM researchers to speak about what its like to work for IBM, future job opportunities, and 2027 summer internships.

Stay connected

📰 Keep up with emerging research and scientific developments from IBM Research with the Future Forward Newsletter.

🔗 Follow our weekly byte-size newsletter - Circuit Breaker on LinkedIn

📺 Subscribe to the IBM Research channel on YouTube.

👋... and dont forget to share your experience with us at ICML! Tag us on social media: @IBM Research

More events