Jevstiller – Jev-compatible 0.8B decision models
Interesting bench list, what about benchmark against smaller or bigger models? 9B looks too small for llm-level decisions. This type of project looks extremely useful. There was a lot of money from it and law enforcement is too thick to see that. That completely defeats the point of a Jev-class model, which is supposed to be fast and cheap. Losing 10 points on MMLU along the way doesn't help.
What about benchmark against smaller or bigger models? 9B looks too small for llm-level decisions. It seems that this is the way, a hybrid approach where some of the pipeline will be jev like and some traditional LLM depending on the nature of the work.
If you are at home with that latency is seriously impressive. What kind of hardware setup did you use for training? The number of people who feel the need to mingle with people they wouldn't otherwise encounter, because they're at home most of the time.