CM
The Career Machine
Back to Jobs
External Redirect Remote

Machine Learning Engineer

Moss 2w ago
Analyzing job security signals...
Employer Website Verification
findwork.dev

About Moss (Wikipedia Record):

Mosses are small, non-vascular flowerless plants in the taxonomic division Bryophyta sensu stricto. Bryophyta sensu lato may also refer to the parent group, bryophytes, which comprises liverworts, mosses, and hornworts. Mosses typically form dense green clumps or mats, often in damp or shady locations. The individual plants are usually composed of simple leaves that are generally only one cell thick, attached to a stem that may be branched or unbranched and has only a limited role in conducting water and nutrients. Although some species have conducting tissues, these are generally poorly developed and structurally different from similar tissue found in vascular plants. Mosses do not have seeds and after fertilisation develop sporophytes with unbranched stalks topped with single capsules containing spores. They are typically 0.2–10 cm (0.1–3.9 in) tall, though some species are much larger. Dawsonia superba, the tallest moss in the world, can grow to 60 cm (24 in) in height. There are approximately 12,000 species.

Platform Portal Preview

Open full site

Moss

findwork.dev

Verified PortalSSL Encrypted
Launch Moss

About the Role

Founding ML Engineer Moss is building the retrieval runtime for real-time AI. We help agents access the right knowledge, conversation history, and user context in milliseconds, and use that context to decide what to do next. We’re looking for a Founding ML Engineer to own the models and machine learning systems behind that experience. You’ll work across embeddings, retrieval, reranking, multilingual understanding, agent intelligence, and our Action Layer; taking ideas from experiments into production. The work comes with real constraints: limited memory, CPU execution, changing context, multiple languages, and latency budgets that leave little room for error. Your job is to improve intelligence and quality while making the models practical to run. What You’ll Do Train and fine-tune embedding and reranking models for real-world retrieval workloads. Build multilingual embedding models that retrieve accurately across languages, regions, and mixed-language conversations. Improve the intelligence behind our Founding Agent. From understanding intent and retrieving context to choosing better responses and converting conversations into meaningful outcomes. Help build Moss’s Action Layer, enabling agents to move from retrieving context to determining and executing the right next action. Own the full model-development cycle: dataset creation, training, evaluation, optimization, deployment, and iteration. Build evaluation pipelines that measure retrieval relevance, multilingual quality, agent outcomes, latency, memory usage, and inference cost. Improve model efficiency through distillation, quantization, and inference optimization, particularly for CPU and ARM devices. Work with runtime and SDK engineers to ship models across cloud, browser, edge, and device environments. Investigate production failure cases and turn them into better datasets, evaluations, and models. Make practical decisions about what to train, what to adapt, and what to ship. Core Stack The work spans: Python and deep learning frameworks for training and experimentation. Embedding models, rerankers, contrastive learning, and semantic retrieval. Multilingual and cross-lingual representation learning. Agent evaluation, intent understanding, tool selection, and action prediction. Dataset curation, hard-negative mining, synthetic data, and reproducible evaluation. Model distillation, quantization, and portable inference. Moss’s Rust runtime and SDKs across cloud and on-device environments. You don’t need to have worked with every part of the stack. You do need to understand how model decisions affect the system running them and the user experience they create. Your First 90 Days From day one: Work directly with our models, evaluation pipelines, Founding Agent, and production use cases. Start contributing code and experiments immediately. By 30 days: Understand the current quality and performance baselines. Own a concrete improvement to a model, dataset, evaluation pipeline, or Founding Agent capability, with evidence that it solves a real problem. By 60 days: Take a model improvement through evaluation and deployment. This could mean improving multilingual retrieval, making the Founding Agent more effective, or advancing an Action Layer capability. Work with the engineering team to validate its behavior under realistic hardware and workload constraints. By 90 days: Independently own a meaningful part of the ML roadmap. Identify the next bottleneck, define the experiments, and drive improvements into production without waiting for a tightly scoped task. What We’re Looking For Experience training or fine-tuning models and deploying them into production. Strong foundations in representation learning, information retrieval, and model evaluation. Strong Python skills and the ability to write maintainable code beyond a research notebook. An understanding of how training data, objectives, and evaluation choices affect real-world model behavior. Ability to reason about tradeoffs between quality, latency, memory, and compute. Comfort working through ambiguous problems and owning the result. Clear communication about what you tried, what worked, what failed, and what should happen next. Nice to Have Experience with embedding models, rerankers, or search relevance. Experience building multilingual or cross-lingual models. Experience evaluating or improving conversational agents. Experience with tool sel

Similar Roles

View all
C
CloudflareHybrid
AIPythonJavaScriptData Quality+1 more
C
CloudflareHybrid or Remote
LeadershipCommunication
Remote
Remote