NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation

LLM-powered multi-agent systems can now automate the full research pipeline from ideation to paper writing, but a fundamental question remains: automation for whom? Researchers operate under different resource configurations, hold different methodological preferences, and target different output formats. A system that produces uniform outputs regardless of these differences will systematically under-serve every individual user, making personalization a precondition for research automation to be genuinely usable. However, achieving it requires three capabilities that current systems lack: accumulating reusable procedural knowledge across projects, retaining user-specific experience across sessions, and internalizing implicit preferences that resist explicit formalization. We propose NanoResearch, a multi-agent framework that addresses these gaps through tri-level co-evolution. A skill bank distills recurring operations into compact procedural rules reusable across projects. A memory module maintains user- and project-specific experience that grounds planning decisions in each user's research history. A label-free policy learning converts free-form feedback into persistent parameter updates of the planner, reshaping subsequent coordination. These three layers co-evolve: reliable skills produce richer memory, richer memory informs better planning, and preference internalization continuously realigns the loop to each user. Extensive experiments demonstrate that NanoResearch delivers substantial gains over state-of-the-art AI research systems, and progressively refines itself to produce better research at lower cost over successive cycles.

AI-assisted analysis from the archive. Read the original sources and publication notes below. Proposed workflows, market estimates and outcome claims need independent verification before a purchasing or operational decision.

From Static Scripts to Evolving Partners: Personalized AI Research with NanoResearch

Executive Summary

Research automation has long struggled with a lack of personalization. Most systems follow rigid templates that fail to account for a scientist's specific resource constraints, stylistic preferences, or historical context. NanoResearch addresses this by introducing a multi-agent framework that uses tri-level co-evolution of skills, memory, and internal policy. By distilling procedural knowledge into reusable skills and learning from implicit user feedback, the system becomes more specialized with every interaction. This approach moves research automation from a generic text-generation tool to a personalized partner that improves output quality while reducing operational costs over time. It represents a significant step toward making autonomous R&D systems viable for diverse enterprise environments.

The Motivation: What Problem Does This Solve?

Existing AI research agents are often designed as one-size-fits-all solutions. They tend to produce uniform outputs that ignore the nuanced requirements of different users. For instance, a researcher at a startup might prioritize low-cost experiments, while a lead at a major laboratory might demand high-fidelity simulations. Current systems lack the ability to accumulate procedural knowledge across projects or retain experience across different sessions. This leads to a repetitive cycle where the user must constantly correct the agent's basic methodological choices. Without personalization, these systems remain academic curiosities rather than reliable productivity tools. Researchers need an agent that adapts to their specific 'way of doing things.'

Key Contributions

  • The introduction of a Skill Bank that distills recurring operations into compact, reusable procedural rules.
  • A dedicated Memory Module that maintains user-specific and project-specific histories to ground planning decisions.
  • A label-free policy learning mechanism that converts free-form feedback into persistent parameter updates.
  • A tri-level co-evolutionary framework where skills, memory, and planning policies reinforce each other over time.

How the Method Works

NanoResearch functions through the continuous interplay of three specialized layers. Unlike traditional agents that start each task from a blank slate, this framework leverages a cumulative knowledge architecture.

Architecture

The system is built on a hierarchy of skill acquisition and application. The Skill Bank acts as a library of 'best practices' learned from previous successes. For example, if an agent successfully optimizes a specific type of neural network architecture, it distills that process into a rule for future use.

Training and Feedback

The most interesting component is the label-free policy learning. Instead of requiring structured data, the system analyzes free-form feedback from the user. It uses this feedback to update the internal planner's parameters, effectively 'internalizing' the user's preferences. This means if a user consistently favors certain statistical tests over others, the system eventually learns to select those tests by default without being explicitly told to do so.

Dataset and Memory

The Memory Module works alongside this policy to ensure context isn't lost between sessions. It stores past successes, failures, and specific project constraints. This grounding ensures that the planner doesn't suggest unrealistic actions that have failed in the past or that contradict the researcher's known resources.

Results: Benchmarks and Performance

While the specific statistical table is not provided in the paper summary, the authors report that NanoResearch delivers substantial gains over current state-of-the-art AI research systems. The framework demonstrates a unique ability to progressively refine itself. Data indicates that as the system completes more cycles, the cost of research decreases because the agent relies on its refined Skill Bank rather than expensive, exploratory planning. The performance improvement is not just in output quality, but in the efficiency of the research process itself.

Strengths: What This Research Achieves

The primary strength of NanoResearch is its longevity and adaptability. It solves the 'reusability' problem by making procedural knowledge a persistent asset. Additionally, the system's ability to internalize implicit preferences is highly promising for enterprise settings where style and protocol are often non-negotiable but rarely documented in a way an AI can naturally understand. It demonstrates high reliability by grounding new plans in historical memory, reducing the likelihood of repetitive errors.

Limitations and Failure Cases

However, the system is not without potential risks. The reliance on implicit feedback could lead to 'echo chambers' where the agent reinforces a user's existing biases or methodological errors. There's also the question of the 'cold-start' phase: how many project cycles are required before the Skill Bank and Policy become truly effective? Furthermore, the computational overhead of maintaining a co-evolving policy might be significant for smaller organizations without high-end infrastructure.

Real-World Implications and Applications

If this technology scales, it could fundamentally change how corporate R&D labs operate. We could see a shift where every lead scientist has a 'digital twin' agent that knows their experimental style and resource limitations. In industries like pharmaceuticals or materials science, this would mean faster iteration on discovery protocols. Engineering workflows would also benefit from agents that can draft documentation and research reports that already align with internal company standards.

Relation to Prior Work

NanoResearch builds on the foundation of multi-agent systems and Large Language Model (LLM) planners. However, it fills a critical gap left by previous state-of-the-art models that treat every interaction as an isolated event. While earlier work focused on the breadth of research tasks, NanoResearch focuses on the depth of the user-agent relationship. It bridges the gap between generic automation and specialized expertise.

Conclusion: Why This Paper Matters

The core insight of this research is that personalization isn't a luxury: it's a requirement for usable AI. By enabling a tri-level co-evolution of skills, memory, and policy, NanoResearch provides a blueprint for agents that actually grow with their users. It proves that a system can become more intelligent and cost-effective through the simple act of doing its job, making it a pivotal development in the field of autonomous research.

Appendix

The full paper and methodology are available at the Hugging Face repository under the paper link provided in the citation. The framework is currently being evaluated for its applicability in automated scientific discovery pipelines.

Recorded source

Hugging Face

The archived text is presented as originally stored. A source link does not mean every statement in the generated analysis is supported by it.

Test a first direction.

Bring your own constraints into the demo. Explore an initial workflow, then discuss the evidence, integrations and evaluation a pilot would need.

Explore this workflow

Browse architecture concepts and decks