Real Estate Sales Jobs in Bahrain
693 Jobs Found
<p>Job Overview We are seeking a dedicated and experienced Assistant Service Manager-Automotive to join our expanding team. With a strong local foundation and a global network, our company is a leader in industries including automotive, healthcare, manufacturing, and real estate. We are committed to excellence and employ state-of-the-art technologies to maintain a competitive edge. The successful candidate will support the Service Manager in overseeing the daily operations of our busy automotive service department, ensuring the highest standards of customer satisfaction, workshop efficiency, and team performance.</p><p>Responsibilities</p><ul><li>Assist the Service Manager in the day-to-day management of the service department and workshop.</li><li>Supervise, mentor, and motivate a team of service advisors and technicians to achieve departmental goals.</li><li>Serve as a key point of contact for customer enquiries, providing exceptional service and resolving any issues or complaints in a professional and timely manner.</li><li>Monitor and control workshop workflow to ensure that all vehicle repairs are completed efficiently and to the highest quality standards.</li><li>Assist in managing workshop loading, scheduling, and job allocation.</li><li>Conduct quality control checks on completed work to ensure it meets our stringent standards.</li><li>Help prepare and analyse service department performance reports.</li><li>Ensure strict adherence to all health and safety regulations within the workshop.</li></ul><p><strong>Desired Candidate Profile</strong></p><ul><li>Proven experience within an automotive service environment, with previous experience in a supervisory or senior technician role.</li><li>Strong technical knowledge of vehicle diagnostics, repair, and maintenance procedures.</li><li>Excellent leadership, communication, and interpersonal skills.</li><li>A strong commitment to delivering outstanding customer service.</li><li>Exceptional organisational and problem-solving abilities.</li><li>The ability to work effectively under pressure in a fast-paced environment.</li><li>Proficiency with workshop management software and other relevant IT systems.</li><li>A full, valid driving licence is essential.</li></ul>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
Senior Director - Investment Risk
<p>Location: Bahrain</p><br>
About the Client
<p>Our client is a leading financial services organisation with a diversified investment platform across multiple asset classes. As the business continues to grow, they are seeking a Senior Director - Investment Risk to support the Chief Risk Officer in strengthening investment risk management, enhancing governance frameworks, and providing independent risk oversight to support strategic investment decisions.</p><br>
Key Responsibilities
<ul>
<li>Lead independent investment risk assessments across a diversified portfolio, including private equity, real estate, private debt, and capital markets investments.</li>
<li>Review investment proposals, due diligence materials, financial models, and valuations to identify key risks and provide independent challenge to investment decisions.</li>
<li>Conduct stress testing, scenario analysis, and risk assessments to evaluate investment exposures, market conditions, and potential downside impacts.</li>
<li>Prepare and present investment risk reporting, insights, and recommendations to senior management and Board-level committees.</li>
<li>Monitor and enhance investment risk frameworks, including risk appetite, policies, limits, and governance processes.</li>
<li>Analyse market conditions and macroeconomic trends to identify emerging risks and assess their impact on investment portfolios.</li>
<li>Partner with investment teams and risk stakeholders while supporting the development and mentoring of junior team members.</li>
</ul>
Key Requirements
<ul>
<li>Bachelor’s degree in Finance, Accounting, Economics, or a related discipline; CFA, MBA, or equivalent professional qualifications are advantageous.</li>
<li>8+ years’ experience within investment risk, investment analysis, private equity, asset management, investment banking, or a related financial services environment.</li>
<li>Strong experience assessing investment opportunities, including due diligence reviews, financial modelling, valuations, and investment risk analysis.</li>
<li>Proven ability to provide independent risk challenge and influence senior stakeholders, investment professionals, and governance committees.</li>
<li>Strong understanding of financial markets, investment products, portfolio risks, and investment risk management methodologies.</li>
<li>Excellent analytical, problem-solving, and communication skills, with the ability to interpret complex information and provide actionable insights.</li>
<li>Commercially minded with strong stakeholder management skills and the ability to operate effectively in a fast-paced investment environment.</li>
</ul><br> </div>
<p><h4>About Apex</h4>
<p>The Apex Group was established in Bermuda in 2003 and is now one of the world's largest fund administration and middle office solutions providers. Our business is unique in its ability to reach globally, service locally and provide cross-jurisdictional services. With our clients at the heart of everything we do, our hard-working team has successfully delivered on an unprecedented growth and transformation journey, and we are now represented by circa 13,000 employees across 112 offices worldwide. Your career with us should reflect your energy and passion. That's why, at Apex Group, we will do more than simply 'empower' you. We will work to supercharge your unique skills and experience. Take the lead and we'll give you the support you need to be at the top of your game. And we offer you the freedom to be a positive disrupter and turn big ideas into bold, industry-changing realities. For our business, for clients, and for you.</p>
<h4>The role & key responsibilities</h4>
<p>The successful applicant will be part of the Apex Fund Administration service offering, performing, and providing fund administration services to real estate funds. The individual will maintain client relationships; take responsibility for coordination of NAV calculation, regulatory reporting, audit preparation, new business take-on and due diligence requirements of the company. The individual will also manage the coordination of regulatory reporting, audit preparation, review and timely delivery of reports. This is an exciting role for someone who is determined to get involved in all aspects of fund administration.</p>
<p><strong>Role responsibilities:</strong></p>
<ul>
<li>Ensuring NAV calculations are reviewed and delivered by the agreed delivery timeframes along with any other client specific reporting requirements</li>
<li>Operational onboarding of new and take on funds</li>
<li>Attendance of meetings where NAV calculations, financial statements and other deliverables are discussed</li>
<li>Calculation and processing of fund capital calls and distributions</li>
<li>Ongoing implementation and refinement of SSAE 18 and ISAE 3402 controls</li>
<li>Ensuring client enquiries are answered in accordance with Apex's service standards on an ongoing basis</li>
<li>Ensuring compliance with regulatory requirements and other requirements of the funds' specifications</li>
<li>Review and authorisation of payments via bank portals</li>
<li>Review of financial statements and related regulatory reports</li>
<li>Ability to deal with auditors (external and internal), manage and respond to audit queries</li>
<li>Monitoring of administration fee payments and chasing of debtors</li>
<li>Ensuring best practices are adopted and improving processes to gain efficiencies</li>
<li>Managing and training staff whilst providing feedback on their performance</li>
<li>Perform investor AML-related risk assessments and due diligence checks</li>
<li>Perform periodic review of investor AML/LYC files</li>
<li>Review FATCA/CRS returns and oversee filing process</li>
<li>Act as the senior point of contact with clients both internal and external</li>
<li>Manage the expectations of clients, seeking and acting on client feedback</li>
<li>Add value to client account planning</li>
<li>Promote new ideas and services by applying knowledge of the industry/sector to create client value</li>
<li>Continually identify improvements and implement efficient ways of working</li>
<li>Create opportunities for cross team collaboration</li>
<li>Act in accordance with the Apex Code of Conduct</li>
<li>Deliver the Apex Client Charter</li>
</ul>
<h4>Skills required</h4>
<ul>
<li>Holds a professional accounting qualification with a strong academic background</li>
<li>At least 10+ years' relevant experience in real estate fund administration role</li>
<li>Knowledge of accounting standards i.e. IFRS, UK, LUX, US GAAP</li>
<li>Ability to work to and meet agreed deadlines</li>
<li>Excellent interpersonal and written communications skills</li>
<li>Ability to multi-task and work in a pressurised environment</li>
<li>Capacity to problem-solve and ability to deal with complex issues</li>
<li>Ability to work independently with minimum supervision/guidance</li>
<li>Must be a self-starter and willing to take initiative</li>
<li>Strong PC skills including Word, Excel and Macros in Excel</li>
<li>Experience in managing people will be considered an asset</li>
<li>Knowledge of eFront will be considered an advantage</li>
<li>Able to present financial and business data to clients adding own analysis and expertise</li>
</ul>
<h4>What you will get in return</h4>
<ul>
<li>A genuinely unique opportunity to be part of an expanding large global business</li>
<li>Working with a strong and dynamic team</li>
<li>Training and development opportunities</li>
<li>Exposure to all aspects of the business, cross-jurisdiction and to working with senior management directly</li>
</ul>
<h4>Additional information</h4>
<p>We are an equal opportunity employer and ensure that no applicant is subject to less favourable treatment on the grounds of gender, gender identity, marital status, race, colour, nationality, ethnicity, age, sexual orientation, socio-economic status, responsibilities for dependants, physical or mental disability. Any hiring decisions are made on the basis of skills, qualifications and experience. We measure our success as a business, not only by delivering great products and services and continually increasing our assets under administration and market share, but also by how we positively impact people, society and the planet.</p></p><p></p>
<section><p class="heading jdMain">Job Description</p><p class="heading">Roles & Responsibilities</p><div class="paragraph"><p>On our Geoscience and Petrotechnical teams, proven expertise and intelligent tech meet, powering our legacy and future of subsurface solutions. Whether in the field or our learning centers, your unique skills and understanding of hydrocarbons will help solve the toughest challenges for clients every day. And with a start at SLB, you ll be set up for a bright future making a real impact across our business and industry. As a Geologist , you will combine your understanding of earth sciences with a comprehensive knowledge of the latest logging technologies to determine reservoir architecture and hydrocarbon potential. You will become adept at multiple software systems and work closely with customers to find innovative ways to solve some of their most complex challenges. As one of our Geophysicists , you will apply your knowledge and expertise of the earth s properties to enhance interpretations of geological data and greater define how we understand the subsurface. We acquire huge amounts of often previously unseen seismic and geophysical data around the world and you will help transform it into the knowledge that powers better decision-making and more effective, more efficient services. You will be involved in the acquisition, processing and interpretation of that data and we offer a range of career opportunities to develop your skills and get exposure across the data lifecycle. As a Petrophysicist , you will combine logging data from multiple downhole sensors with sample data to determine hydrocarbon production capacity, lithology and fluid saturation of the reservoir to ultimately help us optimize its production. You will incorporate data from multiple wells and additional sensors to consider acoustics, spectroscopy and magnetic resonance to enhance overall accuracy or build a clearer picture of the reservoir by understanding its permeability and mechanical properties. As a Reservoir Engineer , you will use data and our leading software products and solutions to create reservoir models that help clients make decisions that deliver safer, optimized, long-term production for each reservoir by simulating fluid flow phase behavior and reservoir physical properties. As a Production Optimization Engineer , you will deliver performance improvements to our clients assets worldwide though virtual representations of our downhole products which incorporate calculations, finite element analysis (FEA), computation fluid dynamics (CFD), costing and parametric modeling into one cohesive system.</p></div></section><section><p class="heading">Desired Candidate Profile</p><p class="paragraph"></p><p>Meet minimum degree requirements</p><p>Technically curious and determined to improve approaches and methods of discovery</p><p>Ambitious and looking to take on responsibility</p><p>Able to effectively contribute to a team</p><p>Good written and verbal communication</p><p>Focused on quality</p><p></p></section>
<p><h4>About Al-Futtaim Group</h4>
<p>Established in the 1930s as a trading business, Al-Futtaim Group today is one of the most diversified and progressive, privately held regional businesses headquartered in Dubai, United Arab Emirates. Structured into five operating divisions; automotive, financial services, real estate, retail and healthcare; employing more than 35,000 employees across more than 20 countries in the Middle East, Asia and Africa, Al-Futtaim Group partners with over 200 of the world's most admired and innovative brands. Al-Futtaim Group’s entrepreneurship and relentless customer focus enables the organization to continue to grow and expand; responding to the changing needs of our customers within the societies in which we operate.</p>
<p>By upholding our values of respect, excellence, collaboration and integrity; Al-Futtaim Group continues to enrich the lives and aspirations of our customers each and every day.</p>
<h4>Overview of the role</h4>
<p>The Beauty Advisor is responsible for advising customers on beauty products and assisting in sales through knowledgeable and engaging interactions. The role involves maintaining high standards of customer service, ensuring product knowledge, and driving sales by promoting featured products and services. The Beauty Advisor also plays a role in merchandising, maintaining stock, and complying with operational procedures to enhance the store's effectiveness and profitability.</p>
<h4>What you will do</h4>
<strong>Sales and customer engagement</strong><br>
<li>Recommend appropriate products and provide information highlighting the features, advantages, and benefits to customers.</li>
<li>Assist customers in locating, selecting, and purchasing desired products based on their preferences.</li>
<li>Utilize selling techniques such as suggestive selling, upselling, and cross-selling to increase Average Transaction Value (ATV).</li>
<li>Attend to all customer queries and needs, ensuring satisfaction and resolving complaints within set standards.</li>
<br><strong>Stock and merchandise management</strong><br>
<li>Monitor availability of GOBE products and update reports on product shelf life.</li>
<li>Ensure proper safekeeping of all store merchandise to prevent shoplifting, damages, and pilferages.</li>
<li>Timely display all products following set planograms and guidelines.</li>
<li>Replenish products on display as needed with complete shelf tag price and price tags.</li>
<br><strong>Operational compliance and standards</strong><br>
<li>Maintain cleanliness and orderliness of the assigned area.</li>
<li>Comply with all set customer service standards, including Smile and Greet and Offers Basket.</li>
<li>Prepare and set up store promotions using marketing collaterals effectively.</li>
<li>Monitor and update price changes and shelf/price tags as needed.</li>
<li>Strictly comply with company policies, including attendance, punctuality, and grooming standards.</li>
<br><strong>Personal development and team collaboration</strong><br>
<li>Accomplish MFP and PDP on time.</li>
<li>Participate in and implement agreed Employee Engagement Survey action plans.</li>
<li>Develop leadership and problem-solving skills.</li>
<li>Proactively contribute with a good team spirit and take initiatives.</li>
<h4>Required skills to be successful</h4>
<li>Customer service excellence and empathy with strong relationship skills.</li>
<li>Proficiency in retail operations and stock management.</li>
<li>Leadership and problem-solving capabilities.</li>
<li>Integrity, trustworthiness, and the ability to deal with ambiguity.</li>
<h4>What qualifies you for the role</h4>
<li>High school diploma or equivalent; degree not essential.</li>
<li>Minimum 2 to 4 years of retail experience with a focus on the beauty or fashion sector.</li>
<li>In-depth knowledge of fashion/beauty industry and trends.</li>
<li>Skills in retail operations, including stock management, visual merchandising, systems, and cash handling.</li>
<li>Excellent customer service, empathy, and results orientation.</li>
<h4>About Al-Futtaim Retail</h4>
<p>Al-Futtaim Retail has established itself as one of the leaders in retail across the Middle East, Africa & Asia over the past 30 years. We have developed partnerships with some of the biggest and most respected brands in the world including IKEA, ACE and Toys R Us in the Middle East and the Inditex Group of Brands (Zara, Mango, Bershka and P&B) across Asia. We are also one of the largest global partners of Marks and Spencer’s in both regions with over 75 stores offering both fashion & food options.</p>
<p>Most recently we have been responsible for bringing brands to the Middle East for the first time with the exciting launches of Watsons and B&Q and we aim to continue to be agile and adaptive to our markets with new launches and further development. For this to be possible we aim to recruit the best talent from all backgrounds who will continue to challenge and develop our diverse workforce which includes over 100 nationalities across 12 countries. Join us today and make a difference…</p></p><p></p>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br>What this opportunity involves: We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT: Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for: 8+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard: Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paidEffort estimate Tasks for this project are estimated to take 30 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation: Up to $200/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~30 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br>What this opportunity involves: We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT: Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for: 8+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard: Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paidEffort estimate Tasks for this project are estimated to take 30 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation: Up to $200/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~30 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves Frontier coding agents are already good at passing tests.<br> We measure whether they pass them the right way .<br>We're building a dataset to evaluate the safety and conduct of AI coding agents — not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.<br>You'll design tasks where the easy path is the unsafe one, and write the tests that catch it: Build realistic developer environments — a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes Write tests that verify whether the agent completed the task the right way — catching corners cut, not just checking outputs Iterate on tasks and tests based on QA feedback — review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT: Not data labeling; Not prompt engineering; Not cybersecurity or red-teaming — there is no attacker in the scenario.<br> Cybersecurity experience is a nice-to-have but not a requirement.<br> We're looking for engineers who understand how code should behave, not penetration testers.<br> Strong software engineers, not security specialists; Not writing code from scratch — the agent writes most of the code; you design the situation and evaluate the outcome; What we look for 4–5+ years in software development; Core stack: Python, JavaScript/TypeScript; Strong test design skills — functional and integration tests that separate safe from unsafe completion, not just correct from incorrect; Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar); Familiarity with GitHub PRs and CI workflows as a user; Stack breadth is welcome, not a filter.<br> Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful — but you don't need to be an expert in every layer; English proficiency — B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> The real difficulty is building the temptation — a scenario where the unsafe or out-of-scope path is the path of least resistance — and then writing tests that reliably catch an agent that took it.<br> Tasks have many valid solutions; tests must accept all of them and reject the bad ones.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Project time expectations For this project, tasks are estimated to require around 20-25 hours per week during active phases, based on project requirements.<br> This is an estimate, not a guaranteed workload, and applies only while the project is active.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation On this project, contributors can earn up to $75 per hour equivalent , depending on their level and pace of contribution.<br> Compensation varies across projects depending on scope, complexity, and required expertise.<br> Please note that other projects on the platform may offer different earning levels based on their requirements.<br></span> </div>
<p><h4>Please submit your CV in English and indicate your level of English proficiency.</h4>
<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.</p>
<h4>What this opportunity involves</h4>
<p>Frontier coding agents are already good at passing tests. We measure whether they pass them the right way. We're building a dataset to evaluate the safety and conduct of AI coding agents — not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.</p>
<p>You'll design tasks where the easy path is the unsafe one, and write the tests that catch it:</p>
<ul>
<li>Build realistic developer environments — a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history</li>
<li>Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes</li>
<li>Write tests that verify whether the agent completed the task the right way — catching corners cut, not just checking outputs</li>
<li>Iterate on tasks and tests based on QA feedback — review agent solutions, analyze failures, and refine until the evaluation is fair and robust</li>
</ul>
<h4>What this is NOT:</h4>
<ul>
<li>Not data labeling;</li>
<li>Not prompt engineering;</li>
<li>Not cybersecurity or red-teaming — there is no attacker in the scenario. Cybersecurity experience is a nice-to-have but not a requirement. We're looking for engineers who understand how code should behave, not penetration testers. Strong software engineers, not security specialists;</li>
<li>Not writing code from scratch — the agent writes most of the code; you design the situation and evaluate the outcome;</li>
</ul>
<h4>What we look for</h4>
<ul>
<li>4–5+ years in software development;</li>
<li>Core stack: Python, JavaScript/TypeScript;</li>
<li>Strong test design skills — functional and integration tests that separate safe from unsafe completion, not just correct from incorrect;</li>
<li>Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar);</li>
<li>Familiarity with GitHub PRs and CI workflows as a user;</li>
<li>Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful — but you don't need to be an expert in every layer;</li>
<li>English proficiency — B2+</li>
</ul>
<h4>Why this is hard</h4>
<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. The real difficulty is building the temptation — a scenario where the unsafe or out-of-scope path is the path of least resistance — and then writing tests that reliably catch an agent that took it. Tasks have many valid solutions; tests must accept all of them and reject the bad ones.</p>
<h4>How it works</h4>
<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid</p>
<h4>Project time expectations</h4>
<p>For this project, tasks are estimated to require around 20-25 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.</p>
<h4>Compensation</h4>
<p>On this project, contributors can earn up to $75 per hour equivalent, depending on their level and pace of contribution.</p>
<p>Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.</p></p><p></p>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves Frontier coding agents are already good at passing tests.<br> We measure whether they pass them the right way .<br>We're building a dataset to evaluate the safety and conduct of AI coding agents — not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.<br>You'll design tasks where the easy path is the unsafe one, and write the tests that catch it: Build realistic developer environments — a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes Write tests that verify whether the agent completed the task the right way — catching corners cut, not just checking outputs Iterate on tasks and tests based on QA feedback — review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT: Not data labeling; Not prompt engineering; Not cybersecurity or red-teaming — there is no attacker in the scenario.<br> Cybersecurity experience is a nice-to-have but not a requirement.<br> We're looking for engineers who understand how code should behave, not penetration testers.<br> Strong software engineers, not security specialists; Not writing code from scratch — the agent writes most of the code; you design the situation and evaluate the outcome; What we look for 4–5+ years in software development; Core stack: Python, JavaScript/TypeScript; Strong test design skills — functional and integration tests that separate safe from unsafe completion, not just correct from incorrect; Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar); Familiarity with GitHub PRs and CI workflows as a user; Stack breadth is welcome, not a filter.<br> Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful — but you don't need to be an expert in every layer; English proficiency — B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> The real difficulty is building the temptation — a scenario where the unsafe or out-of-scope path is the path of least resistance — and then writing tests that reliably catch an agent that took it.<br> Tasks have many valid solutions; tests must accept all of them and reject the bad ones.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Project time expectations For this project, tasks are estimated to require around 20-25 hours per week during active phases, based on project requirements.<br> This is an estimate, not a guaranteed workload, and applies only while the project is active.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation On this project, contributors can earn up to $75 per hour equivalent , depending on their level and pace of contribution.<br> Compensation varies across projects depending on scope, complexity, and required expertise.<br> Please note that other projects on the platform may offer different earning levels based on their requirements.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<p>We help the world run better At SAP, we keep it simple: you bring your best to us, and we'll bring out the best in you. We're builders touching over 20 industries and 80% of global commerce, and we need your unique talents to help shape what's next. The work is challenging but it matters. You'll find a place where you can be yourself, prioritize your wellbeing, and truly belong. What's in it for you? Constant learning, skill growth, great benefits, and a team that wants you to grow and succeed.</p><p>A Solution Advisor works closely with our customers and prospects to identify and solve business challenges and meet their strategic objectives using SAP solutions. As part of the sales team, a Solution Advisor is the subject matter expert responsible for the functional and technical knowledge within the sales cycle. A Solution Advisor provides deal support by participating in discovery sessions, executive meetings and presentations and delivers software demonstrations that help the customer understand SAP s unique value proposition. In addition to deal support, the Solution Advisor participates in marketing events to generate demand, leads Design Thinking sessions, and collaborates with the broader sales team to identify whitespace opportunities at existing accounts.</p><p>As a Solution Advisor within the SAP Next Gen - Academy for Customer Success, you will be responsible to:</p><ul><li>Successfully complete a 10-month program that strengthens the foundation for a successful customer-facing career at SAP.</li><li>Participate in experiential learning opportunities with colleagues from all over the world and acquire a wide variety of business, industry and SAP solution skills while working with emerging and cutting-edge technologies.</li><li>Receive on-the-job training under the mentorship of a senior Solution Advisor colleague while working with our customers to gain real world experience and acquire the skills necessary to help guide our customers through their Digital Transformation journey.</li></ul><p>The program will enrich your knowledge of SAP and give you the professional experience to serve our customers. We offer full-time employment from day one with practical learning applications for your role. After successful completion of the program, you are expected to lead customer discovery sessions and survey activities to uncover business challenges and opportunities for innovation. You will create and deliver high impact and engaging software demonstrations that compel the customers to select SAP over other competitive offerings. You will also provide demand generation support through marketing events and deal execution support by responding to request for proposals.</p><p>The SAP Academy for Customer Success is a global development program designed for talent who are early in their career.</p><p>The SAP Academy for Customer Success offers a three-year journey that drives accountability and enhances productivity. It enables graduates to make a quick impact in customer-facing roles while fostering career longevity and leadership potential.</p><p>Join us for a unique opportunity to build a global network, collaborate with customers to solve real business challenges, and gain hands-on experience with world-class cloud solutions all while learning in a dynamic environment and earning competitive pay and benefits.</p><p>#SAPAcademyforCustomerSuccess</p><p>#SAPNextGen</p><p>SAP Next Gen is our global experience for students and early talent (0 3 years of professional work experience). Being a part of the Next Gen community provides access to a supportive global network, tailored development opportunities, and the ability to contribute to high-impact projects across the business.</p><p>SAP s employees across different regions are enabled to do their best job with the right mix of office and remote work according to country-specific guidelines and regulations. In general, our hybrid work setup consists of three days a week in the office or on-site with customers or partners.</p><p><br></p><p><strong>Desired Candidate Profile</strong></p><p>2-3 years of professional experience with a strong foundation in Finance Topics, technical and business processes, exposure to relevant technologies/solutions, and customer-facing skills.</p><p>Technical and business process knowledge, combined with hands-on experience using relevant technologies and industry-standard tools to support solution delivery and operational efficiency.</p><p>A cooperative and productive approach to working relationships, internally and externally.</p><p>Quote To Cash (Q2C) is strongly preferred</p><p>A strong ability to quickly learn new concepts, adapt to changing environments, and apply knowledge to deliver results.</p><p>An understanding of AI fundamentals, uses, and ethics, to identify business problems solvable with AI.</p><p>Strong Business Acumen, including demonstrated knowledge of business processes and/or industries.</p><p>Proficiency in English to engage with our global network.</p><p>AI focus areas: (AI Adoption Mindset, Agentic AI Day-to Day Practice, and Context Engineering)</p><p>A curious and agile mindset, with a passion for learning and the resilience to adapt in a fast-paced, evolving environment.</p><p>Strong problem-solving, project management, and creative thinking skills, with the ability to collaborate across teams and communicate ideas clearly and effectively.</p><p>Emotional intelligence and cultural awareness that support inclusive teamwork and thoughtful stakeholder engagement.</p><p>A growing foundation in business acumen and emerging technologies, especially artificial intelligence, paired with a results-oriented approach and the courage to take initiative.</p>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br>What this opportunity involves: We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT: Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for: 8+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard: Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paidEffort estimate Tasks for this project are estimated to take 30 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation: Up to $200/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~30 hours each; you set your own schedule.<br></span> </div>
<p><h4>Please submit your CV in English and indicate your level of English proficiency.<\/h4>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves:<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<h4>What this is not:<\/h4>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<h4>What we look for:<\/h4>\n<li>8+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<h4>Why this is hard:<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply? Pass qualification(s)? Join a project? Complete tasks? Get paid.<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 30 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Compensation details are provided upon engagement. Tasks are estimated at approximately 30 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br>What this opportunity involves: We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT: Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for: 8+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard: Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paidEffort estimate Tasks for this project are estimated to take 30 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation: Up to $200/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~30 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> About the Role You’ll design coding tasks that challenge frontier AI coding agents.<br> Each task is a self-contained Docker environment with a broken piece of software; an AI agent attempts the fix; automated tests verify the outcome.<br> Your deliverable is the full task package: broken code, tests, instructions, and a reference solution proving the task is solvable.<br> Responsibilities : Invent a realistic developer scenario — a real bug, a broken ETL, a missing feature — not a toy problem.<br> Build a reproducible Docker environment with pinned dependencies.<br> Write a pytest that verifies outcomes, not specific commands — deterministic, non-flaky, and does not leak the fix.<br> Write an instruction.<br>md that reads like a Jira ticket a developer would receive.<br> Write a reference solve.<br>sh proving the task is solvable.<br> Calibrate difficulty so current state-of-the-art agents solve the task 20–60% of the time.<br> Iterate based on feedback from expert QA reviewers.<br> Later: review other authors’ tasks as a QA reviewer.<br> Not in scope Data labeling, prompt engineering.<br> Production code to ship — you design problems and verification for AI agents.<br> Leetcode puzzles — scenarios must look like real developer work.<br> Not every candidate task ships — quality over quantity.<br> Requirements 3+ years of production software development in one backend stack — Python, Go, Node.<br>js, Java, or Rust.<br> Depth in one stack beats breadth.<br> Python + pytest fluency — required regardless of primary stack.<br> The task harness is pytest-based even when the broken app is in another language.<br> Fixtures, parametrize, monkeypatch, timeouts, conftest.<br>py. Docker authoring — reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user.<br> Linux & Bash — comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail.<br> AI coding agent experience — Claude Code, Cursor, Roo Code, or similar, on non-trivial work.<br> You can cite a specific time the AI was confidently wrong and how you caught it.<br> English — B2+ written.<br> Not a fit Data Science, ML, or Computer Vision engineers without backend-engineering output.<br> Manual QA testers without automation or test authoring.<br> Frontend-only, low-code / no-code, IT Support, or Business Analysts.<br> Engineers who have never written pytest from scratch.<br> Junior, intern, or assistant as the most recent role.<br> Preferred qualifications Domain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (NumPy / PyTorch / SciPy), DevOps, or Git internals.<br> Modern Python tooling (uv, poetry, pyproject.<br>toml). Coverage tooling (pytest-cov, coverage.<br>py, gcov, llvm-cov, kcov).<br> Fuzzing or property-based testing (Hypothesis).<br> Prior contribution to agent-evaluation benchmarks or related frameworks.<br> Process Apply → Pass qualification (90-minute sample-task screen + short behavioral interview) → Join a project → Complete tasks → Get paid.<br> Time commitment Onboarding: ~10 hours per first task.<br> Steady state: ~5 hours per task, 2–4 parallel tasks per author.<br> Realistic weekly load: 8–20 hours.<br> Higher volume available for top performers.<br> You choose when and how to contribute; tasks must be submitted by the deadline and meet acceptance criteria.<br> Compensation: Paid contributions, rates up to $35/hour *.<br> Task-based compensation equivalent to hourly rate, depending on performance and volume.<br> Some projects include incentive payments.<br> *Rates vary based on expertise, skills assessment, location, project needs, and other factors.<br> Higher rates may be provided to highly specialized experts.<br> Lower rates may apply during onboarding or non-core project phases.<br> Payment details are shared per project.<br> Apply Submit your CV via the Mindrift platform.<br> Indicate your English level, note this role (Software Engineering Evaluation Specialist — Terminal Bench), and include a GitHub profile link if available.<br></span> </div>
<p><h4>Please submit your CV in English and indicate your level of English proficiency.<\/h4>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves:<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n <li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n <li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n <li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n <li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is NOT:<\/h4>\n<ul>\n <li>Not data labeling<\/li>\n <li>Not prompt engineering<\/li>\n <li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for:<\/h4>\n<ul>\n <li>8+ years in software development<\/li>\n <li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n <li>Experience writing tests (functional, integration)<\/li>\n <li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard:<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply? Pass qualification(s)? Join a project? Complete tasks? Get paid.<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 30 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation:<\/h4>\n<p>Up to $200\/hr equivalent, depending on level and pace. Tasks are estimated at approximately 30 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description</h4>
<p>Toloka AI supports frontier model post-training by building domain-specific reinforcement learning environments, tasks, and evaluation frameworks designed by real practitioners.</p>
<p>Mindrift, powered by Toloka — a leading enterprise AI and machine learning data partner since 2014 — connects top domain experts with cutting-edge AI initiatives. Backed by Toloka’s deep expertise in scalable data generation, crowd technology, and applied ML systems, Mindrift enables experts to shape how next-generation generative models learn, reason, and perform.</p>
<p>We are launching a management consulting domain focused on translating real-world consulting engagements into structured learning environments for advanced AI systems. To do this credibly, we are assembling a team of strategy consultants from top-tier firms who can convert authentic project experience into end-to-end examples — from problem structuring and work planning to analysis, synthesis, and client-ready recommendations.</p>
<p>You will join a growing team of consultants from leading strategy firms shaping how AI learns high-level business reasoning.</p>
<p><strong>Important:</strong> This role is exclusively for consultants with direct experience at a top-tier strategy consulting firm. If you do not have hands-on project experience at one of the firms listed below, please do not apply. This requirement ensures the domain is built by practitioners trained to the highest standards of structured problem-solving and client delivery.</p>
<p><strong>Eligible firms:</strong> McKinsey & Company, Boston Consulting Group (BCG), Bain & Company, Oliver Wyman, Roland Berger, Monitor Deloitte (Deloitte S&C), EY-Parthenon, Kearney, and Strategy& (PwC).</p>
<h4>Who we’re looking for</h4>
<p>Consultants with 3+ years of experience at one of the firms listed above, with hands-on project experience in:</p>
<ul>
<li>Structuring ambiguous client problems into workable analytical plans</li>
<li>Building financial models, market analyses, or synthesized findings from messy inputs</li>
<li>Producing client-ready deliverables under time pressure</li>
<li>Forming and defending recommendations under uncertainty</li>
</ul>
<p>No deep technical background is required — we will onboard you on the lightweight tools involved.</p>
<h4>What you’ll do</h4>
<ul>
<li>Build realistic consulting project environments — create detailed project scenarios grounded in real engagement dynamics: industry context, financials, constraints, conflicting inputs, and incomplete information.</li>
<li>Design structured consulting tasks for AI agents — break projects into discrete tasks that mirror real consulting work: market sizing, commercial due diligence, cost optimization, growth strategy, operational diagnosis, benchmarking, and more.</li>
<li>Define evaluation criteria and quality standards — develop grading frameworks, evaluation rubrics, and golden-answer solutions for each task, used to train and calibrate an LLM-based grading system that evaluates AI outputs at scale.</li>
</ul>
<p>This is a remote, project-based, individual-contributor role focused on analytical design and evaluation.</p>
<h4>Skills & requirements</h4>
<ul>
<li>3+ years at McKinsey, BCG, Bain, Oliver Wyman, Roland Berger, Monitor Deloitte, EY-Parthenon, Kearney, or Strategy&</li>
<li>Strong structured problem-solving and hypothesis-driven thinking</li>
<li>Ability to translate vague problems into clear analytical steps and deliverables</li>
<li>High attention to logical consistency and output quality</li>
<li>Independent, self-directed working style</li>
<li>Clear written English (B2+)</li>
</ul>
<h4>Compensation</h4>
<p>On this project, contributors can earn up to $60 per hour equivalent, depending on their level and pace of contribution.</p>
<p>Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.</p></p><p></p>