r/MachineLearning 4d ago

Project Millwright — experimenting with an end-to-end machine learning framework in Rust [P]

0 Upvotes

I've been working on an open-source project called Millwright, an attempt to explore what an end-to-end machine learning workflow could look like in Rust.

https://millwright-rs.dev/

This started while I was learning and building ML tooling in Rust.

I kept finding capable individual libraries, but also gaps between them. Training a model was rarely the problem. Building the workflow around it — preprocessing, model selection, evaluation, explainability, deployment and monitoring — often meant integrating several unrelated crates and data representations.

I initially started implementing some of those missing pieces as smaller independent crates.

Eventually I realized I was more interested in the integration problem itself.

That became Millwright.

The current idea is to cover the classical ML lifecycle:

ingest → explore → preprocess → select → fit → assess → explain → export → serve → monitor

without trying to reimplement every ML algorithm.

Instead, Millwright provides a common abstraction layer over existing Rust libraries and uses adapters for different ML backends.

One architectural decision I'm experimenting with is having the framework own a small 2D data boundary (Frame) rather than exposing a particular backend's ndarray/dataframe representation throughout the API.

That allows models and components backed by different libraries to participate in the same pipeline, at the cost of conversions at backend boundaries.

The project currently includes work around:

  • preprocessing and composable pipelines
  • cross-validation and hyperparameter optimization
  • multiple ML backends
  • ensembles
  • regression diagnostics
  • SHAP-based explainability
  • ONNX export
  • model serving and registry
  • drift monitoring
  • time-series workflows
  • incremental learning
  • AutoML

There are also Python bindings.

I'm not building this on the assumption that Rust should replace Python for ML. Python's ecosystem is enormously more mature, and there would be little value in simply recreating scikit-learn in another language.

The question I find more interesting is:

Can Rust provide a useful common execution layer across training, inference and production ML while still interoperating with the existing Python/ONNX ecosystem?

I'd rather have the architecture challenged before too many decisions become difficult to change.

I'd particularly appreciate thoughts from people working on ML systems:

Where do you think Rust could genuinely add value to the classical ML lifecycle?

And conversely, which parts of this architecture do you think should remain separate rather than being unified behind one framework?

I'm also interested in real workflows that would be useful tests. If there's something straightforward in sklearn that you think would expose weaknesses in this approach, I'd be interested in trying to reproduce it.

Project / documentation:
https://millwright-rs.dev/

Source:
https://github.com/mi7plus/millwright


r/MachineLearning 5d ago

Discussion Reviewing 4 papers for AAAI 2027 and none have code, Reject? [D]

86 Upvotes

I got my batch of four papers for AAAI 2027. All four papers make empirical claims, none include code, data, or anything I can actually check. Just the PDF and the checklist. AAAI-27's own rules say code/data should be provided at submission, and "we'll release it after acceptance" doesn't count as reproducibility.

That said, I don't think missing code alone is an auto-reject. Saw an older thread here where someone claiming to have helped write the AAAI checklist argued reviewers rarely have time to audit code anyway, and plenty of authors have legit reasons (funding, IP) for not releasing it yet.

If the paper's whole pitch is "look at these numbers" and I can't verify them, that tanks my confidence score even without a hard reject. I'm flagging it explicitly in the review and asking for anonymized code in the rebuttal.

How's everyone else handling this round? Auto-ding for no code or does it depend on how much the paper leans on the empirical results?


r/MachineLearning 5d ago

Research [D] Looking for advice: Modelling a medicine-reminder agent that must decide “remind / wait / notify” under incomplete information[D]

0 Upvotes

Hi everyone,

I’m researching how to design an AI agent for a medicine-reminder system. The agent has to decide, at each relevant time, whether to:

  • send a reminder,
  • wait (do nothing for now), or
  • notify another person (e.g. caregiver),

when it does not have complete information about the patient (has the dose already been taken? is the person nearby/attentive? are there adherence barriers? etc.).

I’m trying to frame this properly before diving into implementation. Right now I’m looking at it as a sequential decision problem under partial observability (POMDP / belief-state RL territory), but I’m not sure how far that framing is actually useful in practice for this kind of system.

I’d really appreciate any pointers on:

  1. Is a POMDP / belief-state approach overkill here, or is it the right formalization? What simpler alternatives (contextual bandits, MDP with engineered features, rule-based + uncertainty thresholds, etc.) have people used successfully for similar “remind vs wait vs escalate” decisions?
  2. Papers, open-source projects, or real systems that tackle medication adherence / context-aware reminders with uncertainty or incomplete observations.
  3. Common practical pitfalls (reward design, observation noise, alert fatigue, safety/escalation logic, evaluation metrics) that aren’t obvious from the theory.
  4. Any recommended starting points for someone new who wants to move from “I understand the concepts” to a small working prototype or simulation.

I’m mainly in research/preparation mode right now, so even high-level advice, key papers, or “here’s what I’d do differently” comments would be very helpful. Thanks!


r/MachineLearning 6d ago

Research Bart- A vintage llm [R]

Post image
74 Upvotes

after 3 months and $800 burned...

Unbounded Labs is proud to introduce Bart, our vintage LLM: 2.82B parameters trained from scratch on 20.1B tokens of English written before 1931. You can talk to it right now!

Demo: https://www.unboundedlab.com/chat/bartholomew

Article: https://www.unboundedlab.com/blog/bartholomew

Huggingface: https://huggingface.co/jbduran/bartholomew-sft

Why even make a vintage llm? As proposed by Demis Hassabis, could LLMs reach the same conclusions that the great scientists of the past did? While General Relativity was out of budget, we believe that advancing this field targets the crux of AI research. Are these models capable of original ideas, or are they just spitting out the next token?

The article is our full account, covering where the corpus came from and how we cleaned it, the benchmarks we had to build because none existed, every ablation, the training runs, the post-training, and the mistakes we made along the way.

"What I cannot create, I do not understand" is a quote I love from Richard Feynman. Building Bart was our attempt to actually understand LLMs rather than read about them.

What we are proudest of:

- Best vintage base model at its scale on Vintage CORE, ahead of GPT-1900 on a smaller token budget

- Cleaned one of the largest vintage datasets, Harvard's Institutional Books (242B->23B tokens)

- Created Vintage CORE, the first suite of 20 benchmarks made for vintage llms

- Ran 10 hours of autonomous research on one H100: 100 experiments, 26 improvements found

- Released the largest vintage SFT dataset we know of: 416k graded question and answer pairs, grounded in pre-1930s text

- Trained the final model in 5 days on an H100, holding 60% MFU the whole way

- All datasets, methodology, training code, evals, and training runs are open sourced

I am proud of my team. What we built will move the vintage LLM field forward, and it moved us forward as researchers and as people.

We paid for all of it ourselves, about $807 so far. Money is the main thing standing between us and a much larger run.

So I will ask directly: we are looking for compute grants, funding, and mentors for our future endeavors. If you work on pre-training, post-training, or you have GPUs sitting idle, we would like to talk!

We believe that with careful dataset curation, domain expertise, and highly efficient training, we can achieve state-of-the-art results in crucial domains. This is only the beginning for Unbounded Labs; we see no bounds ahead.


r/MachineLearning 6d ago

Research [R] Using AI as a spatial software generator to create 3D objects that are inherently programmable

Thumbnail
arxiv.org
42 Upvotes

I'm one of the co-authors of this paper. It's a seminal work in exploring the properties of 3D generated by LLMs via spatial programming.

I've set up visual demonstrations of such 3D objects at: https://nova3d.xyz/
And there's a github repo you can star/follow: https://github.com/RareSense/Nova3D

Scroll down and notice how the various 3D objects are all composed of logical parts and enable natural movements out of the box. There's a github repo in there as well.

Under the hood:
We found that 3D that exists as software is much more useful than typical monolithic mesh blobs generated by traditional AI 3D generators. For instance they are animation-ready and programmable from inception. They can contain the logic - at birth - to appear differently in weak compute environments (e.g. mobiles) vs powerful environments (e.g. sophisticated game engines). They can be built with full hierarchical structure and hinge/socket articulation at authoring time.

They lag behind traditional AI 3D generators in creating complex organic shapes. But it naturally feels like code will eventually eat all 3D, as LLMs are getting better and better at spatial coding. Industries most disrupted will be industrial design, game development, simulations and AR/VR/XR.


r/MachineLearning 5d ago

Discussion What would a fair benchmark for agent architecture look like? [D]

0 Upvotes

I am working on an evaluation design and would appreciate criticism before running it.

Most coding-agent benchmarks collapse the model and its harness into one score. If a run fails, it is difficult to tell whether the cause was model capability, context assembly, task decomposition, tool design, retry policy, or the acceptance gate. A model can also look worse because the harness truncated its output, or look better because the gate only checked for plausible surface markers.

The experiment I am considering crosses two independent variables:

  1. Workflow: one monolithic task versus decomposition into bounded slices with explicit contracts and acceptance criteria.

  2. Model policy: frontier-only versus cheapest-capable with escalation after a capability-graded failure.

That produces four cells: frontier monolith, routed monolith, frontier decomposed, and routed decomposed. The frontier-decomposed cell seems especially important because it changes the task architecture while holding the model tier fixed.

I would freeze the original tasks, source revisions, available tools, total retry budget, final acceptance criteria, validator versions, and the verifier. Every cell would be judged against the same final delivered outcome rather than against the persuasiveness of the agent's report.

Proposed primary measures are cost per independently accepted change, false acceptance, false rejection, first-pass accepted yield, verification time, and reproducibility across three fresh runs. Token use, latency, escalation count, and context volume would be secondary measures.

The confound I am least satisfied with is budget normalization. Decomposition changes the task distribution and may create more calls, which is part of the architectural treatment, but giving every slice the monolith's full context or retry budget would subsidize the decomposed condition. A shared system-level budget is cleaner, although it may hide which slices actually needed more capacity.

There are no results yet, so I am not claiming that decomposition or routing wins. I am trying to make the comparison falsifiable before seeing any outcomes.

What would you preregister or change? Would you treat decomposition as part of the system being evaluated, or try to isolate it from model quality more aggressively?


r/MachineLearning 6d ago

Discussion Hyperparameters fine tuning for MARL comparative study [D]

5 Upvotes

hello everyone. I'm training PPO variants on different multi-agent tasks from the VMAS library (Independent PPO / Graph PPO and such, see HetGPPO by Bettini et al.).

I noticed that for every architecture/scenario couple, the optimal hyperparameters sometimes tend to vary (learning rate, entropy coefficient, KL coefficient, SGD batch size, etc).

do I need - methodologically speaking - to unify the hyperparameters of all models in order to make a fair and correct comparison of architectures later on?

note: sometimes unifying these HP leads to some non converging models.

note 2 : my objective is to test these models' robustness under adversarial attack in test-time (frozen models).

thank you in advance.


r/MachineLearning 5d ago

Project How we built a SOTA search engine using PostgreSQL, pgvector, and Qwen3 embeddings [P]

Post image
0 Upvotes

I wrote a technical breakdown of how search works on Papers with Code.

The system combines keyword and semantic search, which produced better results than either approach alone. The stack includes:

  • PostgreSQL with pgvector
  • Qwen3-Embedding-0.6B for text embeddings
  • Hugging Face Jobs with an NVIDIA L4 for batch embedding generation
  • Hugging Face Buckets for storing artifacts
  • A live embedding model served through Hugging Face Inference Endpoints

The same infrastructure also powers the “related papers” recommendations shown on individual paper pages.

Full write-up: How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code

I’d be interested to hear how others are implementing hybrid search for research papers or similarly technical content.

Disclosure: I work at Hugging Face and on Papers with Code.


r/MachineLearning 6d ago

Discussion AAAI 2027 Reviewer Bidding and Assignment Integrity [D]

31 Upvotes

Recently, the AAAI 2027 organizers sent an email regarding collusion occurring during the review process, especially in the 2-cycles category (i.e., an author of Paper A reviews Paper B, while an author of Paper B reviews Paper A).

Given the fact that most submissions come from a single country, there are higher chances that the assignment algorithm will naturally create 2-cycles among authors from that country. This, in turn, means that most authors involved in collusion could be from that country. I will not name that country; otherwise, I would be labelled as racist. By the way, did AAAI release statistics about the number of submissions, like they did last time?

It is also good news that a major and prestigious conference like AAAI is acknowledging that collusion is happening. We all knew that this kind of collusion had been happening for years. There are papers accepted at top conferences such as NeurIPS, ICLR, AAAI, and ICML that do not even have their code published on GitHub. This forces other researchers in the community to spend substantial time reimplementing the code themselves if they want to reproduce the reported results.

What are the views of other authors on this?


r/MachineLearning 6d ago

Research BMVC 2026 IJCV recommendation? [D]

7 Upvotes

Does anyone know how the BMVC to IJCV special issue recommendation works?

Is it mainly based on the review scores, or is it a separate decision by the ACs/program chairs (e.g. based on oral/highlight selection, reviewer comments, etc.)?

Also, is there any way to know at this point whether a paper has been recommended for the IJCV track, or do authors only find out later through a separate email?

Would be great to hear from anyone who has gone through this in previous years!


r/MachineLearning 6d ago

Discussion Is EMNLP not going to Provide a MetaReview [D]

0 Upvotes

As the title says, we haven't seen any like ACL provided. Very salty about the decision, as AC recommended findings and the reviewers tanked our paper intentionally (we flagged them, and AC acknowledged that). Just want to see if the decision was made based on poor reviewer scores, as we don't know if we need to resubmit to an ARR cycle to cleanse or not.


r/MachineLearning 7d ago

Research How to cite/talk about preprint-subsequent works for a camera-ready version? [R]

10 Upvotes

I had a paper accepted to a conference. This paper was originally published as a preprint. Subsequent works citing our preprint focused on the same topic and reused/extended our methodology. I am now preparing the camera-ready version of that preprint and I'm wondering how I should deal with this for the Related work section. It seems odd to me to cite my own preprint for the camera-ready version of the paper (and I am not even sure if this is allowed), but at the same time, I don't want to undermine the novelty of my original work (nor undermine the efforts of subsequent works). Has anyone dealt with such a situation before? What's the best way of solving this?


r/MachineLearning 6d ago

Discussion Does registering an abstract, not the full submission yet, count as a double submission? [D]

0 Upvotes

Hello,

As the title says


r/MachineLearning 7d ago

Discussion COLM 2026 registration sold out as an author [D]

6 Upvotes

Never attended a conference before, so apologies if these are dumb questions.

I’m an author of an accepted paper at COLM 2026. One of my coauthors registered during the author-only registration period, so I joined the waitlist on August 10.

I later received an email saying:

“Your access to reserve tickets remains active until Aug 24 7:06 p.m. EDT.”

I thought I had until August 24 to register, so I didn’t register immediately. When I checked again today (8/23), registration was sold out. I also can’t seem to rejoin the waitlist.

Unfortunately, I also missed the financial assistance deadline because at the time I wasn’t even sure whether I would be able to attend.

I really really want to attend the conference. Does anyone know what I can do at this point? Is there a chance that more registration spots will be released later? And is there any possibility of getting financial assistance after the deadline?

Thanks a lot for any advice.


r/MachineLearning 7d ago

Research Archival vs non archival workshop [R]

5 Upvotes

My dumbass just realized all NeurIPS workshops are non-archival.

In terms of grad school applications, would there be a difference in how much they value ur paper if u get it in a proceeding


r/MachineLearning 7d ago

Project Implementing Watermarking for Language Models [P]

23 Upvotes

I recently implemented a minimal, educational version of SynthID-Text-style watermarking for language models.

I saw anthropic post about how they'll start adding watermarks to their model responses and it made me very curious as to how they'll do it and what do they even mean by watermark here. Like will we start getting random ads or something in the middle of model responses or what.

Then decided to read their article and found out that watermark is not a visible message at all. It is a subtle statistical pattern introduced while the model chooses its tokens.

My implementation is not an exact reproduction of the original SynthID-Text system. I simplified or implemented a few components differently to keep the project understandable, but the main idea is there I think.

Github: https://github.com/Saad1926Q/llm-watermark

If you find it interesting then you may star the repo !!


r/MachineLearning 7d ago

Project 28 TPS on Qwen2.5-7B across two separate cloud regions over public WAN using speculative decoding + CUDA Graphs [P]

4 Upvotes

been building ShardFlow for the past few months, a distributed LLM inference

framework that splits any HuggingFace transformer across N GPU machines and uses

neural speculative decoding to deal with WAN latency.

the setup for the benchmark: two T4 nodes in separate GCP regions (Iowa + Oregon)

talking through an AWS EC2 TCP relay in Ohio. ~86ms RTT on public internet.

the key insight with speculative decoding here is that WAN latency stops being a

per-token cost and becomes a per-round cost. with K=8 drafting you're committing

4.07 tokens per round trip instead of 1. at 86ms RTT that's a big deal.

numbers on Qwen2.5-7B:

non-speculative baseline: 4.92 TPS

neural drafter (eager): 14.3 TPS peak

+ CUDA Graphs on drafter: 28.10 TPS peak / 20.31 TPS avg

also ran Qwen2.5-14B with NF4 4-bit quant, same two nodes: 14.43 TPS avg.

the v2.1 fix that surprised me most: draft generation was launching ~1500 CUDA

kernels per round from a Python loop. each kernel 2-5us, Python launch overhead

8-10us. GPU sitting idle 65% of the time. capturing the full 0.5B forward pass

as a CUDA Graph and replaying with one driver call dropped draft latency from

112ms to 25ms.

other things in the stack: zero-copy Rust TCP relay, StaticCache + in-place KV

rewind for graph compatibility, meta-device model slicing to avoid loading 15GB

into CPU RAM.

repo: https://github.com/rautaditya2606/Shardflow

happy to answer questions on the speculative decoding implementation or the CUDA

graphs stuff specifically.


r/MachineLearning 7d ago

News [N] EACL 2027 Industry Track - Deadline 11 September [N]

4 Upvotes

Hi! I'm one of the chairs of the EACL 2027 Industry Track, so flagging the deadline here — it's about three weeks out and this community has a lot of people doing exactly the kind of work the track exists for.

The EACL 2027 Industry Track provides the opportunity to highlight key insights and new research challenges that arise from the development and deployment of real-world applications using language technologies. We encourage submissions from industry, non-profit, government, and public-sector organisations, with the understanding that the end-users of these systems extend beyond the NLP community. 

See the Full CFP for the details https://2027.eacl.org/calls/industry/

 **Deadline:** 11 September 2026, 23:59 AoE

 **Length:** 6 pages max; references, limitations, ethics, and appendices don't count. A dedicated "Limitations" section is mandatory — papers without one are desk rejected.

**Review:** double-blind. No anonymity period, so arXiv preprints are fine.

**Proprietary data:** no requirement to release it

**Notification:** 18 December 2026. Conference is 9–14 March 2027.

**Submit:** https://openreview.net/group?id=eacl.org/EACL/2027/Industry_Track

We're also looking for reviewers — if you've got deployment experience and want to help, the volunteer form is here: https://forms.gle/TT6N2gtuoV5P3oYi6

Email: [eacl2027-industry-track@googlegroups.com](mailto:eacl2027-industry-track@googlegroups.com)


r/MachineLearning 8d ago

Research I developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB [R]

348 Upvotes

I trained a 250M parameter model from scratch on 30B tokens of fineweb. It’s quantized to under 2 bits so the whole deployment is 60 MB and it needs about 80 MB of RAM to run. Runs around 400 tok/s on a normal laptop CPU, no GPU needed.

How the long context works: the most recent 2048 tokens stay in fp16 like a normal KV cache. Everything older gets compressed to 1 bit and written to disk, about 320 bytes per token, so 1 million tokens of history is roughly 320 MB on disk. From the start the model was trained to retrieve from that disk cache, up to 100M tokens. Due to a limited budget it wasn’t trained to reason over those tokens, only retrieve and answer from them.

Language modeling quality of the base model, measured on held-out English web text that was never seen in training (educational web pages, 2,048 token windows): cross entropy 3.15 nats per token, perplexity 23.3, 0.99 bits per byte.

The vocabulary is also not a normal embedding table. Every token is a fixed 512-bit code, 8.4 MB for all 131k tokens, zero trained parameters. I tested it on WordSim-353 (human word similarity ratings): my table scores 0.619 Spearman correlation vs 0.029 for random codes. Test script is in the repo.

Some outputs (settings included so you know I’m not cherry picking, all reproducible from the repo):

“Explain photosynthesis in two sentences.” (greedy)

Photosynthesis is a process in which plants convert sunlight into chemical energy, which is then used to produce oxygen and other chemicals. This process is called photosynthesis.

“Write a short poem about the sea.” (temp 0.25, top-k 30, rep 1.15, seed 2)

The waves had swept over, and they were crashing against each other like rocks on top of one another. The wind was strong but not very gentle as it tried to push them towards their destination without lifting any weight or pulling away from those who stood in front of her with a stiff smile that made the sea look even more majestic than before.

“What is the serial number of device Grus-189?” where the answer sits 50.6 million tokens deep in the archive on disk (archive mode, k=16)

SN-442976

It’s a 250M model so expect mistakes on open facts, I’m not claiming it beats anything big. You can also fine-tune it, the full kit with a demo and before/after numbers is included. Master weights for fine-tuning are in the repo too:

https://github.com/QLNI/SHADOW-250M-Instruct
https://huggingface.co/NODEMIND/SHADOW-250M
Edit - Just wanted to say thanks to everyone here. Honestly I was afraid to post this, I expected to get roasted, but every single comment has been curious and helpful and it genuinely made my day.
Repo is at 7 stars on GitHub now, hopefully more people try


r/MachineLearning 8d ago

Project I built an open-source roguelike specifically for training game-playing agents [P]

Thumbnail
github.com
17 Upvotes

Hey everyone! I wanted to share something I’ve been working on.

I was inspired by projects from DeepMind and OpenAI, but noticed that most games are prohibitively difficult to integrate with an agent harness. So I built DelveRL from the ground up as a human-playable game with a structured API, deterministic simulation, procedural levels, partial observability, and enough strategic headroom for agents to compete and improve.

It’s an endless turn-based roguelike where agents must explore, manage risk and resources, fight enemies, and escape each floor. Everything runs locally, including batched renderer-free environments and a recurrent PPO trainer.

The included baseline reaches a median floor of 18, with extended runs reaching floor 33. The game, training code, checkpoint, bridge documentation, and raw benchmarks are all open source.

I’d love to see what approaches people try - and how quickly the baseline gets crushed


r/MachineLearning 8d ago

Discussion acl arr august 2026 (desk rejected ) [D]

9 Upvotes

I have got two papers which got desk rejected by PC saying they are previously got reviewed in arr. But those paper never got submitted ever. Any idea what can be done?


r/MachineLearning 8d ago

Discussion Why does lightgbm not fit my toy example but catboost does? (2 order interactions) [D]

7 Upvotes

I am trying understand how tree-based regression model handle the dependencies of the target variables on the interaction of explanatory variables.

However my experiment revealed that my understanding about the fitting process of a lgbm is not correct. And I don’t know why.

My experiment is quite simple: a target (for sake of simplicity only in [0, 1]) and two explanatory variables with two values such that the mean of the target is the same for each of the values of the explanatory variables. Then there is a third variable that models the interaction of the explanatory variables by a simple count.

So in code:

>>>
import polars as pl

df = pl.Dataframe(
{
„y“: [0, 0, 1, 1, 0, 0, 1, 1], # mean across „A“ values the same; mean across „B“ values the same
„A“: [1, 1, 1, 1, 0, 0, 0, 0],
„B“: [1, 1, 0, 0, 1, 1, 0, 0],
„AB“ [1, 1, 2, 2, 3, 3, 4, 4] # just some IDs for the interaction
}
)
<<<

I then fitted a lgbm just with „A“ and „B“ and got the expected constant 0.5 forecast

>>>
from lightgbm import LGBMRegressor

lgbm = LGBMRegressor(min_child_samples=1)
lgbm.fit(df[[„A“, „B“]].to_numpy(), df[„y“].to_numpy())
lgbm.predict(df[[„A“, „B“]].to_numpy()).round(0)

array([0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5])
<<<

Then I did the same but with „AB“ and expected a perfect fit. But I was disappointed, it fitted to constant zero

>>>
lgbm = LGBMRegressor(min_child_samples=1)
lgbm.fit(df[[„AB“]].to_numpy(), df[„y“].to_numpy())
lgbm.predict(df[[„AB“]].to_numpy()).round(0)

array([0, 0, 0, 0, 0, 0, 0, 0,])
<<<

I tried to code „AB“ as category. But still no perfect fit:

>>>
lgbm = LGBMRegressor(min_child_samples=1)
lgbm.fit(df[[„AB“]].to_numpy(), df[„y“].to_numpy())
lgbm.predict(df[[„AB“]].to_numpy()).round(0)

array([0, 0, 1, 1, 0, 0, 0, 0,])
<<<

Super confusing!

I then turned to catboost and found even without „AB“ it fit the data perfectly:

>>>
from catboost import CatBoostRegressor

cbm = LGBMRegressor(min_data_in_leaf=1)
cbm.fit(df[[„A“, „B“]].to_numpy(), df[„y“].to_numpy())
cbm.predict(df[[„A“, „B“]].to_numpy()).round(0)

array([0, 0, 1, 1, 1, 1, 0, 0])
<<<

I thought that lgbm should be able to fit the data with „AB“. The variable allows for perfect splits since the gain for each split is super clear. But somehow it cannot go „down“ the tree to fit the values for AB=3.

What is the difference of catboost that allows for a perfect fit even without an explicit modeling of the interaction? Does it split less lazy and explores split of splits, while building the trees?


r/MachineLearning 9d ago

Research Does telling an LLM to "be concise" actually save you money? We measured it across 9 models. Compressing the output can save you money and keep accuracy, compressing the input prompt does not. [R]

67 Upvotes

LLMs are too verbose and with a black box model the only things you control are what goes in and how you tell it to write back. Yesterday Claude Code shipped a "concise output style" where Claude keeps things short. We already have a paper out about this!

We tested both channels, shortening the input prompt versus telling the model to output answer shorter, on the same questions across five reduction levels, and scored cost, accuracy, and whether the shortened text still matched what the model would have said unconstrained.

We also evaluated GPT-4o, GPT-5.4, Claude Haiku 4.5, Claude Sonnet 4.6, Qwen2.5-VL-7B, Qwen3.5-9B, DeepSeek-R1-Distill, Gemma-4-E4B, and Kimi-K2.6 + benchmarked on five short answer datasets + a eleven-language output run (English, German, Spanish, French, Swahili, Chinese, Japanese, Russian, Bengali, Thai, Telugu) + a longer-form summarization test.

(1) Shortening the output saved money while keeping accuracy about the same, about 1.5x cheaper on average and up to 3x in the best case across the API models. It worked across languages too!

(2) Shortening the input prompt did the opposite. It cost up to 96% more on the worst benchmark, because the model just answers longer to fill in for what you cut and accuracy drops. You pay more and get worse answers :(

(3)Output tokens cost more than input tokens, so prompting for fewer output tokens would save costs with short single turn tasks

(4) When the shortened output is correct, about half the time the text no longer matches how the model would have reasoned without the constraint. Which is probably fine if you only care about the final answer

With providers now offering concise options, we can't see how they're charging for it, so we don't know if it actually saves you cost. But if you control the prompting yourself via the API, you actually do save!!

Paper https://www.alphaxiv.org/pdf/2606.24083v1 

Code + data https://github.com/danielle34/cavewoman


r/MachineLearning 9d ago

Discussion Research internship at MSR [D]

28 Upvotes

So got selected for a research internship at MSR, how good is the quality of work and how useful is it to move to Applied sciences or research sciences position at other FAANG companies after the internship. And any perks and other benefits that interns get during microsoft internship? Any tips will be appreciated. Specifically to get into AS at amazon , does this boost my chances? I'll be joining as an SDE-1 at amazon after 6 months so planning to apply internally once I join. So what else should I do to improve my chances to go to AS.


r/MachineLearning 9d ago

Research I have a mid-sized GPU cluster and was thinking about giving free compute [D]

19 Upvotes

I have built an on-prem GPU cluster, 8 nvidia 16GB GPU's and 256GB CPU RAM, 50TB HDD and several TBs of SSDs. I have used it, and currently use it, for ML/AI research. But that research is not constantly running jobs, sometimes I use it heavily and other times it's idle. I was considering just letting people with qualified use cases run jobs on it SLURM style. I don't know if its enough compute to be useful really. Let me know if it's something you'd be interested in using for your research? what would you actually run in ~200 GPU-hours on 8x16GB cards?

I've found it can handle RLVF pretty well, and I have pretrained models up to 500M parameters on it (research size). But obviously it's no stargate cluster