Dataset and Qwen2.5-3B adapters (SFT → GRPO) for multi-turn mystery investigation with tools, unlock DAGs, and cited case closes.
Vaidik
VaidikML0508
AI & ML interests
exploring another way to use gradient decent
Organizations
None yet
Shark Tank Deal Evaluator
This collection features Llama-3.2-3B models fine-tuned to simulate Shark Tank deal evaluations and decision-making based on company pitches and offer
-
VaidikML0508/Shark-Tank-Offer-Evaluator-llama3.2-3B-Instruct-GRPO-16bits-V1
Text Generation • 3B • Updated • 13 • 1 -
VaidikML0508/Shark-Tank-Offer-Evaluator-llama3.2-3B-Instruct-SFT-DPO-4bits-V1
Text Generation • 3B • Updated • 16 -
VaidikML0508/SharkTank-Offer-V1
Viewer • Updated • 255 • 10 -
VaidikML0508/SharkTank-Offer-DPO-dataset-V1
Viewer • Updated • 263 • 8 • 1
Mystery Investigation Agent
Dataset and Qwen2.5-3B adapters (SFT → GRPO) for multi-turn mystery investigation with tools, unlock DAGs, and cited case closes.
Shark Tank Deal Evaluator
This collection features Llama-3.2-3B models fine-tuned to simulate Shark Tank deal evaluations and decision-making based on company pitches and offer
-
VaidikML0508/Shark-Tank-Offer-Evaluator-llama3.2-3B-Instruct-GRPO-16bits-V1
Text Generation • 3B • Updated • 13 • 1 -
VaidikML0508/Shark-Tank-Offer-Evaluator-llama3.2-3B-Instruct-SFT-DPO-4bits-V1
Text Generation • 3B • Updated • 16 -
VaidikML0508/SharkTank-Offer-V1
Viewer • Updated • 255 • 10 -
VaidikML0508/SharkTank-Offer-DPO-dataset-V1
Viewer • Updated • 263 • 8 • 1
small pretraining dataset