<thead id="dkgrx"><rt id="dkgrx"><dfn id="dkgrx"></dfn></rt></thead>

<noscript id="dkgrx"><span id="dkgrx"></span></noscript>

溫馨提示×

溫馨提示×

您好，登錄后才能下訂單哦！

密碼登錄×

忘記密碼？

登錄注冊×

獲取短信驗證碼

其他方式登錄

點擊登錄注冊即表示同意《億速云用戶服務(wù)條款》

用戶登錄×

賬戶密碼登錄

請使用微信掃描上方二維碼

使用幫助

請求超時！

請點擊重新獲取二維碼

hadoop如何通過cachefile來避免數(shù)據(jù)傾斜

發(fā)布時間：2021-12-09 16:26:19 來源：億速云閱讀：225 作者：小新欄目：大數(shù)據(jù)

這篇文章主要介紹了hadoop如何通過cachefile來避免數(shù)據(jù)傾斜，具有一定借鑒價值，感興趣的朋友可以參考下，希望大家閱讀完這篇文章之后大有收獲，下面讓小編帶著大家一起了解一下。

package hello_hadoop;
import java.io.BufferedReader;
import java.io.BufferedWriter;
import java.io.File;
import java.io.FileInputStream;
import java.io.FileReader;
import java.io.FileWriter;
import java.io.IOException;
import java.io.InputStream;
import java.io.InputStreamReader;
import java.net.URI;
import java.net.URISyntaxException;
import org.apache.commons.logging.Log;
import org.apache.commons.logging.LogFactory;
import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.filecache.DistributedCache;
import org.apache.hadoop.fs.Path;
import org.apache.hadoop.io.LongWritable;
import org.apache.hadoop.io.Text;
import org.apache.hadoop.mapreduce.Job;
import org.apache.hadoop.mapreduce.Mapper;
import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;
import org.apache.hadoop.mapreduce.lib.input.FileSplit;
import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;
public class GetFileName {
	//得到處理的文件名，以及將需要的文件緩存到相應(yīng)的節(jié)點
	private final static  Log LOG = LogFactory.getLog(GetFileName.class);
	public static void main(String[] args) throws IOException, ClassNotFoundException, InterruptedException, URISyntaxException {
		LOG.info("Go into the main method ......");
		Configuration conf = new Configuration();
		Job job = Job.getInstance(conf);
		
		job.setJarByClass(GetFileName.class);
		job.setMapOutputKeyClass(Text.class);
		job.setMapOutputValueClass(Text.class);
		job.setMapperClass(GetFileNameMapper.class);
		job.setNumReduceTasks(0);
		//集群上添加DistributedCache data134:9000namenode的描述   #thelinkofthefile改文件的鏈接，下面讀取的時候需要使用
		job.addCacheFile(new URI("hdfs://data134:9000/home/tmp.txt#thelinkofthefile"));
		FileInputFormat.addInputPath(job, new Path(args[0]));
		FileOutputFormat.setOutputPath(job, new Path(args[1]));
		boolean test = job.waitForCompletion(true);
		LOG.info("End  the main method ......");
		System.exit(test?0:1);		
	}
}
class GetFileNameMapper extends Mapper<LongWritable, Text, Text, Text>{
	private final Log LOG = LogFactory.getLog(GetFileNameMapper.class);
	
	@Override
	protected void setup(Mapper<LongWritable, Text, Text, Text>.Context context)
			throws IOException, InterruptedException {
		if(context.getCacheFiles().length>0);
		URI u = context.getCacheFiles()[0];
		//這里使用鏈接來訪問文件
		BufferedReader br = new BufferedReader( new FileReader(new File("./thelinkofthefile")));
		String line = br.readLine();
		context.write(new Text(line), new Text());
		System.out.println("Here I read Line :"+line);
	}
	@Override
	protected void map(LongWritable key, Text value, Mapper<LongWritable, Text, Text, Text>.Context context)
			throws IOException, InterruptedException {
	}
}

感謝你能夠認真閱讀完這篇文章，希望小編分享的“hadoop如何通過cachefile來避免數(shù)據(jù)傾斜”這篇文章對大家有幫助，同時也希望大家多多支持億速云，關(guān)注億速云行業(yè)資訊頻道，更多相關(guān)知識等著你來學(xué)習(xí)!

向AI問一下細節(jié)

推薦閱讀：

免責(zé)聲明：本站發(fā)布的內(nèi)容（圖片、視頻和文字）以原創(chuàng)、轉(zhuǎn)載和分享為主，文章觀點不代表本網(wǎng)站立場，如果涉及侵權(quán)請聯(lián)系站長郵箱：is@yisu.com進行舉報，并提供相關(guān)證據(jù)，一經(jīng)查實，將立刻刪除涉嫌侵權(quán)內(nèi)容。

上一篇新聞：
怎么用Elasticsearch打造知識庫檢索系統(tǒng)
下一篇新聞：
hadoop中mapreduce如何實現(xiàn)串聯(lián)執(zhí)行

猜你喜歡

AI
助
手

產(chǎn)品服務(wù)

地區(qū)劃分

專題活動

幫助支持

關(guān)于我們

售后咨詢

7*24小時在線電話：400-100-2938

7*24小時在線 QQ：800811969

關(guān)注億速云

億速云公眾號

手機網(wǎng)站二維碼

<wbr id="orbpm"><center id="orbpm"><u id="orbpm"></u></center></wbr>

<tr id="orbpm"></tr>

<legend id="orbpm"><font id="orbpm"></font></legend>

<address id="orbpm"></address>

<pre id="orbpm"><form id="orbpm"></form></pre>